Detailed Action
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Amendments
This action is in response to amendments filed December 5th, 2025, in which Claims 1, 3, 7, 9, 14, & 17 have been amended. No claims have been added and claims 2, 8, 10-12, & 18 have been cancelled. The amendments have been entered, and Claims 1, 3-7, 9, & 13-17 are currently pending.
Response to Arguments
Regarding the applicant’s traversal of the 112(f) interpretations of the previous office action, the applicant’s arguments filed December 5th, 2025 have been fully considered, and are unpersuasive.
Applicant asserts that the claims have been amended to recite that the units are implemented using a processor and a memory executing specific instructions, thereby providing sufficient structure to perform the recited functions.
The examiner respectfully submits that the amendments describe An apparatus having at least one processor and a memory including instructions which when executed, result in the units and their functionalities. The actual structure of each unit is still unclarified and still recites the language which invokes 112(f) as before. Therefore, the 112(f) interpretations are maintained.
Regarding the applicant’s traversal of the 112(b) rejections of the previous office action, the applicant’s arguments filed December 5th, 2025 have been fully considered, and are persuasive. The examiner notes that all suggested amendments and clarifications with this regard were made, so therefore, the rejections under 35 U.S.C. 112(b) have been withdrawn.
Regarding the applicant’s traversal of the 35 U.S.C. 101 rejections of the previous office action, the applicant’s arguments filed December 5th, 2025 have been fully considered, and are unpersuasive.
Applicant asserts that the amended limitation, “…An apparatus for refining data and improving a performance of a behavior recognition model by reflecting time-series characteristics of a behavior, the apparatus comprising: at least one processor; and memory including instructions, that when executed by the at least one processor, result in…” is evidence of integration into a practical application because the claimed embodiments use a processor and a memory to execute instructions to perform receiving and outputting data, generate refined dataset, and analyze and synchronize similarities between datasets, which are more than mental processes. Further, applicant clarifies that the recitation of non-generic structural elements such as a processor and memory, and by featuring features such as real-time data, sensor data, etc., a concrete and practical improvement in behavior recognition models is achieved.
The examiner respectfully submits that, as shown in the MPEP 2106, the recitation of the use of a computer to perform an abstract idea is not indicative of integration into a practical application, and neither is the recitation of data gathering or output. The examiner thus fails to see what the “concrete and practical improvement in behavior recognition models” is that applicant is referring to, and may better respond if this is clarified further. Respectfully, the rejections under 35 U.S.C. 101 are thus, maintained.
Regarding the applicant’s traversal of the 35 U.S.C. 102/103 rejections of the previous office action, the applicant’s arguments filed December 5th, 2025 have been fully considered, and are unpersuasive.
Applicant asserts that by amending claims 1 and 14 to include limitations previously cited from now cancelled claims 2, 8, & 10-12, that the cited prior art now fails to teach all the limitations as described.
The examiner respectfully admits that doing this has resulted in WU & KIM, as previously relied upon to reject claims 1 and 14, no longer teach all of the limitations but that by bringing the other previously cited references into the combination, the claimed invention can be taught and that sufficient motivation to combine them exists, as written more fully in the rejections below. Therefore, the rejections under 35 U.S.C. 103 have been maintained.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “a data pre-processing unit configured to receive training data and real-time data as input, identify a missing value of sensor data, and interpolate the sensor data”, “a behavior recognition unit configured to, through a behavior recognition model, generate a behavior recognition classification result for the preprocessed real-time data”, “a data refinement unit configured to correct the behavior recognition classification result to generate a refined dataset”, “a learning model update unit configured to analyze a similarity of the refined dataset and, based on a result of the analysis, perform learning to generate the behavior recognition model”, and “an information output unit configured to express a corrected behavior recognition result to a user” in claim 1.
Further limitations include “a data receiving unit configured to receive the training data and the real-time data used for behavior recognition from various sensors”, “a missingness identification unit configured to identify whether an error or missing value exists in the received real-time data”, and “a data interpolation unit configured to, when it is identified by the missingness identification unit that newly input real-time data needs to be corrected, search for samples having a pattern similar to a pattern of the input real-time data in a database for sample data, infer and generate a value corresponding to an error or missing value of the input real-time data on the basis of sample data having a highest similarity, and interpolate the input real-time data using the generated value” in claim 1.
Further, limitations include “a recognition result correction unit configured to correct a result of the behavior recognition model by reflecting time-series characteristics of a behavior that appears consecutively in a range of a movement of a human body”, “a refined data set generation unit configured to divide the sensor data corresponding to a corrected recognition result (a label) to generate the refined dataset including a pair of [label, sensor data]”, and “a representative pattern data generation unit configured to generate representative pattern data identified as a representative value of the behavior in the refined dataset and store the representative pattern data in a database for representative pattern sample data” in claim 1.
Further, limitations include “a dataset similarity analysis unit configured to analyze a similarity of a dataset used for learning” and “a behavior recognition model generator configured to analyze a similarity of a dataset and, on the basis of the result of the analysis, perform learning to generate the behavior recognition model” in claim 1.
Further limitations include “a recognition model synchronization unit configured to train the behavior recognition model using the training data including a behavior label (a correct answer sheet) and a sensor dataset corresponding to the behavior label” in claim 7.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-18 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea (mental process) without significantly more.
Regarding claim 1, in Step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “An apparatus for refining data and improving a performance of a behavior recognition model”. An apparatus is one of the four statutory categories of invention.
In Step 2a Prong 1 of the 101-analysis set forth in the MPEP 2106, the examiner has determined that the following limitations recite a process that, under the broadest reasonable interpretation, covers a mental process but for recitation of generic computer components:
“identify a missing value of sensor data, and interpolate the sensor data” (A person can mentally evaluate sensor data and make a judgement to identify a missing value and interpolate it (MPEP 2106). For example, if we see that the outside temperature was 89 degrees at 12pm, and 92 degrees at 2pm, we can identify a missing value at 1pm and interpolate/estimate it to be about 90-91 degrees, based on nearby data.)
“…through a behavior recognition model, generate a behavior recognition classification result for the preprocessed real-time data” (A person can mentally evaluate real-time data using their own learned “behavior recognition model” and make a judgement to “classify” that data by recognizing behavior (MPEP 2106). For example, seeing a person running (real-time data) and because we know what running is and how it looks based on its features (behavior recognition model), we mentally classify what we see as a person running.)
“analyze a similarity of the refined dataset and, based on a result of the analysis, perform learning to generate the behavior recognition model” (A person can mentally evaluate the refined dataset to analyze similarity and make a judgement to learn from it, and generate a “model” to recognize it (MPEP 2106).)
“identify whether an error or missing value exists in the received real-time data” (A person can mentally evaluate the received data and make a judgement to identify whether an error or missing value exists (MPEP 2106).)
“infer and generate a value corresponding to an error or missing value of the input real-time data on the basis of sample data having a highest similarity, and interpolate the input real-time data using the generated value” (A person can mentally evaluate a the most similar sample data to a missing value and judgement to infer and generate a value based on that, and interpolate or estimate the input data using that value (MPEP 2106).)
“divide the sensor data corresponding to a corrected recognition result to generate the refined dataset including a pair of a label and the sensor data” (A person can mentally evaluate the sensor data corresponding to a label and make a judgement to divide it based on that label to generate a refined dataset (MPEP 2106).)
“generate representative pattern data identified as a representative value of the behavior in the refined dataset” (A person can mentally evaluate the behavior of the refined dataset and make a judgement to generate representative pattern data / a value that represents the behavior (MPEP 2106).)
“analyze a similarity of a dataset used for learning” (A person can mentally evaluate the similarity of a dataset used for learning (MPEP 2106).)
“analyze a similarity of a dataset and, on the basis of the result of the analysis, perform learning to generate the behavior recognition model” (A person can mentally evaluate the refined dataset to analyze similarity and make a judgement to learn from it, and generate a “model” to recognize it (MPEP 2106).)
“performs the learning only when a similarity between a dataset having previously been used and the refined dataset is less than or equal to a threshold” (A person can mentally evaluate a similarity between a dataset and its refined dataset and make a judgement to only learn from it if it is significantly different (MPEP 2106).)
If claim limitations, under their broadest reasonable interpretation, covers performance of the limitations as a mental process but for the recitation of generic computer components, then it falls within the mental process grouping of abstract ideas. According, the claim “recites” an abstract idea.
In Step 2a Prong 2 of the 101-analysis set forth in MPEP 2106, the examiner has
determined that the following additional elements do not integrate this judicial exception into a
practical application:
“An apparatus for refining data and improving a performance of a behavior recognition model by reflecting time-series characteristics of a behavior” (Uses a computer as a tool to perform an abstract idea (MPEP 2106.05(f)).)
“a data pre-processing unit configured to receive training data and real-time data as input” (Adding insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g)).)
“a behavior recognition unit configured to…” (Mere instructions to apply the judicial exception (MPEP 2106.05(f)).)
“a data refinement unit configured to correct the behavior recognition classification result to generate a refined dataset” (Mere instructions to apply the judicial exception (MPEP 2106.05(f)).)
“a learning model update unit configured to…” (Mere instructions to apply the judicial exception (MPEP 2106.05(f)).)
“an information output unit configured to express a corrected behavior recognition result to a user” (Adding insignificant extra-solution activity (mere data output) to the judicial exception (MPEP 2106.05(g)).)
“a data receiving unit configured to receive the training data and the real-time data used for behavior recognition from various sensors” (Adding insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g)).)
“a missingness identification unit configured to…” (Mere instructions to apply the judicial exception (MPEP 2106.05(f)).)
“a data interpolation unit configured to, when it is identified by the missingness identification unit that newly input real-time data needs to be corrected, search for samples having a pattern similar to a pattern of the input real-time data in a database for sample data” (Adding insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g)).)
“wherein the data refinement unit includes: a recognition result correction unit configured to correct a result of the behavior recognition model by reflecting time-series characteristics of a behavior that appears consecutively in a range of a movement of a human body” (Mere instructions to apply the judicial exception (MPEP 2106.05(f)).)
“a refined data set generation unit configured to…” (Mere instructions to apply the judicial exception (MPEP 2106.05(f)).)
“a representative pattern data generation unit configured to…” (Mere instructions to apply the judicial exception (MPEP 2106.05(f)).)
“store the representative pattern data in a database for representative pattern sample data” (Adding insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g)).)
“wherein the learning model update unit includes: a dataset similarity analysis unit configured to…” (Mere instructions to apply the judicial exception (MPEP 2106.05(f)).)
“a behavior recognition model generator configured to…” (Mere instructions to apply the judicial exception (MPEP 2106.05(f)).)
“wherein the behavior recognition model is configured to, through learning and optimization being performed with a new refined dataset, have various parameters (a weight, a bias, etc.), a layer having a learnable parameter (a convolutional layer, a linear layer, etc.), a value of a registered buffer, and an optimizer and hyperparameter thereof changed to reflect the new refined dataset” (Generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h)).)
“wherein the learning model update unit…” (Mere instructions to apply the judicial exception (MPEP 2106.05(f)).)
Since the claim as a whole, looking at the additional elements individually and in combination, does not contain any other additional elements that are indicative of integration into a practical application, the claim is “directed” to an abstract idea.
In Step 2b of the 101-analysis set forth in the 2019 PEG, the examiner has determined that the claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception.
As discussed above, additional element (xi) recites use of a computer as a tool to perform the abstract idea, which is not indicative of significantly more. Additional elements (xii), (xvi), (xvii), & (xix) recite insignificant extra-solution activities. Further, elements (xii), (xvii), & (xix) recite steps of receiving/transmitting data via a network, which has been determined by the courts to recite a well-understood, routine, and conventional activity, which is not indicative of significantly more (Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362). Further, element (xvi) recites steps that present output of data, which the courts have found to be a well-understood, routine, and conventional activity, which is not indicative of significantly more (Presenting offers and gathering statistics, OIP Techs., 788 F.3d at 1362-63, 115 USPQ2d at 1092-93).). Further, element (xxiii) recites steps that store and retrieve information in memory to be a well-understood, routine, and conventional activity, which is not indicative of significantly more (Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015)). Additional elements (xiii), (xiv), (xv), (xviii), (xx), (xxi), (xxii), (xxiv), (xxv), & (xxvii) recite mere instructions to apply the judicial exception, which is not indicative of significantly more. Additional element (xxvi) recites generally linking the use of the judicial exception to a particular technological environment or field of use, which is not indicative of significantly more. Considering the additional elements individually and in combination, and the claim as a whole, the additional elements do not provide significantly more than the abstract idea. Therefore, the claim is not patent eligible.
Regarding claim 3, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 3 recites “wherein the sensor includes one or more of an accelerometer sensor, a gyroscope sensor, a geomagnetic sensor, an electrocardiogram sensor, a heart rate sensor, a respiration sensor, a skin temperature sensor, and a skin conductivity sensor” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Regarding claim 4, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 4 recites “wherein the data pre-processing unit is configured to, for use in pre-processing, manage a database for representative pattern sample data including samples of various representative values matching a behavior targeted for recognition among datasets having been used for training the behavior recognition model” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Regarding claim 5, it is dependent upon claim 4, and thereby incorporates the limitations of, and corresponding analysis applied to claim 4. Further, claim 5 recites “wherein the data pre-processing unit is configured to, in response to a behavior that is repeatedly observed in training datasets or identified to be eligible to have a meaning in addition to the behavior targeted for recognition, add data corresponding to the behavior to the database for the representative pattern sample data and manage the database” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Regarding claim 6, it is dependent upon claim 4, and thereby incorporates the limitations of, and corresponding analysis applied to claim 4. Further, claim 6 recites “wherein the data pre-processing unit is configured to: when data having been used for the learning includes various users or sensors, manage generalized or general-purpose representative pattern data in a database; and when data having been used for the learning includes data from a specific user or specific sensor, manage personalized or dedicated representative pattern data in a database.” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Regarding claim 7, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 7 recites “wherein the behavior recognition unit further includes a recognition model synchronization unit configured to train the behavior recognition model using the training data including a behavior label (a correct answer sheet) and a sensor dataset corresponding to the behavior label, and the behavior recognition unit is repeatedly updated using information about the behavior recognition model received from the learning model update unit later so as to be kept synchronized with latest information” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Regarding claim 9, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 9 recites “wherein the time-series characteristics of the behavior includes one of. a constraint of an order of transition between behaviors (sequentiality of behaviors) and a causal necessity of transition between behaviors; a transition time between behaviors taking into account a reaction time of a human body and a duration of a behavior (continuity of a behavior); and a movement having a chance of repetitively occurring (periodicity of a behavior)” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Regarding claim 13, it is dependent upon claim 1, and thereby incorporates the limitations of, and corresponding analysis applied to claim 1. Further, claim 13 recites “wherein the learning model update unit repeats a process of synchronizing with the behavior recognition model of the behavior recognition unit until the similarity between the datasets converges” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Regarding claim 14, in Step 1 of the 101-analysis set forth in MPEP 2106, the claim recites “A method of refining data and improving a performance of a behavior recognition model”. A method is one of the four statutory categories of invention.
Further, claim 14 recites similar additional limitations as claim 1, and is rejected under the same rationale, with the following addition: “A method of refining data and improving a performance of a behavior recognition model by reflecting time-series characteristics of a behavior” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.)
Regarding claim 15, it is dependent upon claim 14, and thereby incorporates the limitations of, and corresponding analysis applied to claim 14. Further, claim 15 recites “wherein the generating of the behavior recognition model includes repeating a process of synchronizing with the behavior recognition model of the behavior recognition unit until the similarity between the datasets converges” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Regarding claim 16, it is dependent upon claim 14, and thereby incorporates the limitations of, and corresponding analysis applied to claim 14. Further, claim 16 recites the following additional mental process:
“analyzing the similarity on the basis of a dataset having previously been used” (A person can mentally evaluate the similarity on the basis of a previously used dataset (MPEP 2106).)
Further, claim 16 recites “wherein the generating of the behavior recognition model includes: in the analyzing of the similarity of the dataset, receiving a newly generated refined dataset as input” (In step 2A, prong 2, this recites insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g).) In step 2B, the courts have found steps of transmitting/receiving data over a network to be a well-understood, routine, and conventional activity (Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362).
Further, claim 16 recites “when the similarity is less than or equal to a threshold, returning the refined dataset received as input so that the refined dataset is used for recognition model training” (In step 2A, prong 2, this recites insignificant extra-solution activity (mere data gathering) to the judicial exception (MPEP 2106.05(g).) In step 2B, the courts have found steps of transmitting/receiving data over a network to be a well-understood, routine, and conventional activity (Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362).
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Regarding claim 17, it is dependent upon claim 16, and thereby incorporates the limitations of, and corresponding analysis applied to claim 16. Further, claim 17 recites “wherein the generating of the behavior recognition model includes: additionally using training data having not been used for learning to verify a performance of the behavior recognition model; or tuning various parameters having been used in the model to prevent overfitting of behavior recognition model; and obtaining high performance even for new data not shown in learning” (In step 2a, prong 2, this recites generally linking the use of the judicial exception to a particular technological environment or field of use (MPEP 2106.05(h).) In step 2B, generally linking the use of the judicial exception to a particular technological environment or field of use is not indicative of significantly more.)
Since the claim does not recite additional elements that either integrate the judicial exception into a practical application, nor provide significantly more than the judicial exception, the claim is not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 3-5, 7, 9, & 13-17 are rejected under 35 U.S.C. 103 as being unpatentable over Wu, T. et al. “CARMUS: Towards a General Framework for Continuous Activity Recognition with Missing Values on Smartphones.” Available on June 22 2018 (hereafter, WU), and further in view of PCT Patent Application KR 101969450 B1, published on April 16 2019 (hereafter, KR450), Mojorad, R. et al. “Automatic Classification Error Detection and Correction for Robust Human Activity Recognition.” Available on January 30 2020 (hereafter, MOJORAD), Thornton, C. et al. “Auto-WEKA: Combined Selection and Hyperparameter Optimization of Classification Algorithms.” Available on August 18 2012 (hereafter, THORNTON), & Bifet, A. et al. “Learning from Time-Changing Data with Adaptive Windowing.” Available in 2007 (hereafter, BIFET)
Regarding claim 1, WU teaches “An apparatus for refining data and improving a performance of a behavior recognition model by reflecting time-series characteristics of a behavior”:
([Abstract] “This paper presents the CARMUS framework for continuous activity recognition (behavior recognition of real-time data) with missing values on smartphones (an apparatus). Besides the power and resource constraints discussed in existing work, our framework is proposed to further tackle the critical issue of missing values during data collection. We demonstrate the issue's impact on continuous recognition through a motivating example, and specify two challenges-blackouts and resource constraints-with respect to smartphone-based sensing and processing platforms. To address the challenges, CARMUS provides a novel framework which involves a light-weight admission control unit and a data imputation unit intuited by the daily repeated pattern and temporal smoothness of human activity data (refining the results by reflecting time-series characteristics such as “the daily repeated pattern” or “temporal smoothness of human activity data”). Based on extensive experiments conducted on a real-world data set with 37% of the data missing, we show that the CARMUS framework is effective for achieving an 85.5% recognition accuracy by adopting the state-of-the-art imputation algorithms. (improving a performance of behavior recognition)”)
And further:
([Introduction, B. Proposed Approach & Contributions, Paragraph 2] “The proposed CARMUS framework tackles the blackouts issue by exploring the daily repeated pattern of activity data inspired by the daily routine of human activities [15] and the circadian rhythms [16]. More specifically, we obtain a data matrix by aligning different days' data with the same time of the day (time series data and characteristics), and estimate the missing values by referring to data collected in previous days in this matrix (refining data by reflecting time-series characteristics of a behavior). Through this technique, imputation algorithms designed for multi-dimensional time series [13], [14] can naturally fit into our framework to process data collected in multiple days.”)
Further, WU teaches “a data pre-processing unit configured to receive training data and real-time data as input”:
([Introduction A1, Data Collection, Paragraph 1] “One male subject is asked to carry two Android-based smartphones with data collection service installed—Phone G and Phone D (The phones equivalate to a data pre-processing unit receiving the data), where Phone G is for ground-truth collection (training data) which is required to be always carried by the subject during the day; and Phone D is for real-life data collection following the subject's daily usage patterns2 (real-time data). During data collection, we focus on the missing values caused by the subject not carrying the phone and omit the power constraint by using additional power sources.”)
Further, WU teaches “identify a missing value of sensor data, and interpolate the sensor data”:
([Introduction A1, Data Collection, Paragraph 1] “One male subject is asked to carry two Android-based smartphones with data collection service installed—Phone G and Phone D (data pre-processing unit receiving the data), where Phone G is for ground-truth collection (training data) which is required to be always carried by the subject during the day; and Phone D is for real-life data collection (real-time data) following the subject's daily usage patterns2. During data collection, we focus on the missing values caused by the subject not carrying the phone (identification of missing values) and omit the power constraint by using additional power sources.”)
Fig. 1 illustrates the distribution of effective data collected by Phone D over three days. The effectiveness of the data is determined by the sensor data admission control unit introduced later in Sec III-B. As shown in the figure, despite that the collection service is always kept alive, we still experience an average 55.9% of data loss over time because the subject is not carrying the phone (identification of missing data), which is consistent with the results presented in [11]. Note the above experiments are conducted with the subject aware of the collection process, and the results are not intended to represent a general pattern of missing values for daily smartphone usage for different subjects. Instead, we focus on studying the impact of missing values on the performance of continuous activity recognition by using the subject's data as a case study.
And further:
([Introduction, B. Proposed Approach & Contributions, Paragraph 2] “The proposed CARMUS framework tackles the blackouts issue by exploring the daily repeated pattern of activity data inspired by the daily routine of human activities [15] and the circadian rhythms [16]. More specifically, we obtain a data matrix by aligning different days' data with the same time of the day (time series data and characteristics), and estimate the missing values by referring to data collected in previous days in this matrix (interpolating/estimating the missing values of the sensor data base on nearby data in the time-series). Through this technique, imputation algorithms designed for multi-dimensional time series [13], [14] can naturally fit into our framework to process data collected in multiple days.”)
Further, WU teaches “a behavior recognition unit configured to, through a behavior recognition model, generate a behavior recognition classification result for the preprocessed real-time data”:
([Section III. CARMUS Design, B. Sensor Data Admission Control] “We first introduce the design of the sensor data admission control unit which is responsible for identifying ineffective data and preventing them from polluting the data stream fed to the activity recognition component. (pre-processing the real-time data)
1) Problem Description and Proposed Solution
For activity recognition tasks discussed in this paper, the inertial sensor data are effective only when the user is carrying the phone. As a result, we model the problem of determining the effectiveness of sensor data as a binary classification problem with the two classes corresponding to the user carrying and not carrying the phone, respectively. More specifically, we model the problem as follows.
Given the sensor data collected in a time window of l samples starting at time t, Dt,l=<dt,dt+1,…,dt+l−1 >, assign a label to Dt,l to be carried or not carried by the user, where dt is the sensor data collected at time t.
To solve the above problem, we propose the solution based on two observations: 1) the phone is often left unmoved when the user is not carrying it; 2) in most cases, the phone is left on a surface such as a table with its screen facing up or down. (behavior recognition) The first observation indicates that the variance of acceleration readings is an effective indicator. And the second observation leads us to focus on the z-axis to compare its readings with 1G ≈ 9.8m/s2. As a result, we propose to use the mean and variance of the z-axis readings in Dt,l to represent the features of readings contained in the window.
Fig. 3 illustrates the distribution of instances with the user carrying / not carrying the phone in the 2D feature space of the the z-axis's mean and variance with a window size of l=1s and sampling rate of 50Hz. As shown by the figure, instances belong to different classes discriminates well with each other in the feature space as expected. We then conduct an experiment by applying a J48 decision tree classifier. The resulting detection accuracy with ten-fold-cross-validation is 97.9%, suggesting the effectiveness of the proposed solution. Here, a classification result is attained through the behavior recognition based on pre-processed real-time data”
Further, WU teaches “a data refinement unit configured to correct the behavior recognition classification result to generate a refined dataset”:
([Section III. CARMUS Design, B. Sensor Data Admission Control, Last Paragraph] “Since the ineffective data are not useful for the later tasks, it is desirable to keep the data collection and processing costs to a minimum when the user is not carrying the phone. More specifically, the admission control unit can duty-cycle the data effectiveness detection task by activating the sensors periodically and hibernate on detecting the data are ineffective. While the detailed scheduling strategy is application specific, one important question is what is the minimum amount of data necessary to make an accurate detection for the effectiveness of the data? The above proposed solution has already limited the sensor to be the z-axis of acceleration based on our observations, it is beneficial to further determine the minimal sampling rate r and window size l for lower costs, which we obtain in the next section (further refinement after the previous classification result) through benchmarking”)
And further:
([Section III. CARMUS Design, C. Missing Value Imputation, Paragraph 1] “The raw data stream passed the above sensor data admission control unit (to generate the behavior recognition classification result) only contains effective data that represent the user's activity. However, as shown by the motivating example in Sec. I-A, it is infeasible to obtain reliable continuous recognition results by only using the effective data. As a result, it is important to complete the data stream to obtain reasonable estimations of the missing values. (Further refining the dataset)”)
And further:
([Section III. CARMUS Design, C2. Proposed Solution] “…we design the imputation framework in CARMUS following the nature of daily human activities:
Daily Repeated Patterns. Based on existing evidences of daily routine of human activities [15] and human circadian rhythms [16], we propose to tackle the blackouts by correlating the current missing values with non-missing values collected (or estimated) at the same time of the day in previous days.
Temporal Smoothness. Existing work on sensor-based activity recognition suggests the similarity of sensing values over time [17] for human activities. Following this observation, data collected in the near past are more correlated/similar to the current values. As a result, we propose to address the resource constraints issue by limiting the volume of data processed by the imputation algorithm to only include the data in the closest past.”)
These two steps show how the data is being further refined.
Further, WU teaches “wherein the data pre-processing unit comprises: a data receiving unit configured to receive the training data and the real-time data used for behavior recognition from various sensors”
([Introduction A1, Data Collection, Paragraph 1] “One male subject is asked to carry two Android-based smartphones with data collection service installed—Phone G and Phone D (The phones equivalate to a data pre-processing unit receiving the data), where Phone G is for ground-truth collection (training data) which is required to be always carried by the subject during the day; and Phone D is for real-life data collection (real-time data) following the subject's daily usage patterns2. During data collection, we focus on the missing values caused by the subject not carrying the phone and omit the power constraint by using additional power sources.”)
And further:
([Section III. Carmus Design, A. Overview, Paragraph 1] “…To perform activity recognition on the smartphones, CARMUS collects raw data from three inertial sensors: acceleration (with gravity), linear acceleration (without gravity), and angular velocity from the gyroscope (This citation simply shows that data is collected from “various sensors”). The activation and sampling rate of the inertial sensors are controlled by the Sensor Control Unit according to the control feedback from the sensor data admission control unit.”)
Further, WU teaches “a missingness identification unit configured to identify whether an error or missing value exists in the received real-time data”:
([Introduction A1, Data Collection, Paragraph 1] “One male subject is asked to carry two Android-based smartphones with data collection service installed—Phone G and Phone D (data pre-processing unit receiving the data), where Phone G is for ground-truth collection (training data) which is required to be always carried by the subject during the day; and Phone D is for real-life data collection (real-time data) following the subject's daily usage patterns2. During data collection, we focus on the missing values caused by the subject not carrying the phone (identification of missing values) and omit the power constraint by using additional power sources.”)
Fig. 1 illustrates the distribution of effective data collected by Phone D over three days. The effectiveness of the data is determined by the sensor data admission control unit introduced later in Sec III-B. As shown in the figure, despite that the collection service is always kept alive, we still experience an average 55.9% of data loss over time because the subject is not carrying the phone (identification of missing data), which is consistent with the results presented in [11]. Note the above experiments are conducted with the subject aware of the collection process, and the results are not intended to represent a general pattern of missing values for daily smartphone usage for different subjects. Instead, we focus on studying the impact of missing values on the performance of continuous activity recognition by using the subject's data as a case study.
Further, WU teaches “a data interpolation unit configured to, when it is identified by the missingness identification unit that newly input real-time data needs to be corrected, search for samples having a pattern similar to a pattern of the input real-time data in a database for sample data, infer and generate a value corresponding to an error or missing value of the input real-time data on the basis of sample data having a highest similarity, and interpolate the input real-time data using the generated value”
([Section III. CARMUS Design, C2. Proposed Solution] “…we design the imputation framework in CARMUS following the nature of daily human activities:
Daily Repeated Patterns. Based on existing evidences of daily routine of human activities [15] and human circadian rhythms [16], we propose to tackle the blackouts by correlating the current missing values with non-missing values collected (or estimated) at the same time of the day in previous days (correlating the missing values with other values in a database for sample data).
Temporal Smoothness. Existing work on sensor-based activity recognition suggests the similarity of sensing values over time [17] for human activities (searching for values on the basis of being similar). Following this observation, data collected in the near past are more correlated/similar (similar based on a pattern) to the current values. As a result, we propose to address the resource constraints issue by limiting the volume of data processed by the imputation algorithm to only include the data in the closest past (searching a specific sample of data in a database based on similarity).”)
Following the above analysis, we design the missing value imputation unit in our CARMUS framework as illustrated in Fig. 5. More specifically, the missing data imputation process is composed of three steps.
Matrix Segmentation
Given the current data involving missing values, we first obtain a data matrix by segmenting the current and history data buffered in the system as shown in the dashed box in Fig. 5(a). Data collected in different days are aligned by the time of the day according to the daily repeated activity pattern (corresponding to a pattern) discussed above. To explore the temporal smoothness among the data, we also include part of the last observed data into the segment (a similar pattern).
As shown by the dashed box in Fig. 5(a), the size and position of the data matrix are determined by three parameters—the number of days d, the segmentation size s, and the overlapping rate λ which is the portion of effective data (or estimated previously) in the current segment. Note the CARMUS framework is designed to perform online imputation, which means only the observed data will be included in the current segment. For applications that do not have an online requirement, it is trivial to extend the current framework to include both the observed and future data into the current segment.
Initialization
Before applying the missing value imputation algorithm introduced next, we first initialize the missing values by averaging the last observed data and data with the same time of the day in previous days as shown in Fig. 5(b) (infer and generate a value corresponding to an error or missing value of the input real-time data on the basis of sample data having a highest similarity). In [13], the authors use linear interpolation to initialize the missing values (and interpolate the input real-time data using the generated value). In this paper, we show by our experiment results that the proposed initialization approach improves the system's performance compared to linear interpolation in Sec. V-D2.
Missing Value Imputation
As shown in Fig. 5(c), after the above preprocessing steps, the last step is to complete the data segment by missing value imputation. In this work, we adapt the DynaMMo [13] algorithm for this task because it explores both the temporal smoothness and spatial correlations among different sensors, which well fits into our scenario. It is important to note that the main focus of this paper is to propose a general framework for missing value-tolerant activity recognition which can integrate different imputation algorithms such as DCMF, Non-negative Matrix Factorization, etc. And we leave the topic of choosing and designing more effective algorithms for our future work.
To be more specific, we adapt DynaMMo [13] into our framework with the following steps. Data collected with the same time of the day i across d days form the observation at time i, xi (data at time i for each day contain readings from the three sensors). Each xi corresponds to a hidden state zi through a projection matrix G with an additional Gaussian noise term ϵi∼N(0,Σ), i.e.,
PNG
media_image1.png
27
558
media_image1.png
Greyscale
where the projection matrix G models the spatial correlation among different axes across different days.
The hidden state at the next time step is temporally correlated to the hidden state at the current step through a transition matrix F and an additional Gaussian noise term ωi∼N(0,Λ), i.e.,
PNG
media_image2.png
23
552
media_image2.png
Greyscale
where the transition matrix F models the temporal smoothness among the data [13].
The model parameters are estimated by maximizing the log-likelihood of the joint distribution of hidden states and the observations through Expectation-Maximization. And the missing values are estimated using the resulting series of hidden states. (infer and generate a value corresponding to an error or missing value of the input real-time data on the basis of sample data having a highest similarity) We omit further details in this paper because we focus on the design of the framework instead of the imputation algorithms. Readers interested in imputation algorithms can refer to [13], [14], [18] for details.
Further, WU teaches “wherein the data refinement unit includes: a recognition result correction unit configured to correct a result of the behavior recognition model by reflecting time-series characteristics of a behavior that appears consecutively in a range of a movement of a human body”:
([Page 851, Col. 2, Paragraph 5] “The proposed CARMUS framework tackles the blackouts issue (missing data to be corrected via interpolation) by exploring the daily repeated pattern of activity data inspired by the daily routine of human activities [15] and the circadian rhythms [16] (reflecting time series data). More specifically, we obtain a data matrix by aligning different days’ data with the same time of the day, and estimate the missing values by referring to data collected in previous days in this matrix (reflecting time-series characteristics of a behavior reflecting time-series characteristics of a behavior that appears consecutively). Through this technique, imputation algorithms designed for multi-dimensional time series [13], [14] can naturally fit into our framework to process data collected in multiple days. Experiment results suggest by increasing the number of days from 1 to 4, the activity recognition accuracy increases from 83.9% to 90.6% on the benchmarking data set, showing the effectiveness of our design.”)
And further:
([Page 851, Col. 1, Paragraph 3] “2) Preliminary Results: To understand the issue’s impact on continuous activity recognition performance, we evaluate the recognition accuracy using Phone D’s data (sensor data) by comparing the recognized activities against the results obtained by Phone G (ground-truth sensor data) which is always kept on the subject’s body. We consider three activities including idle, walking, and biking (activities based upon a range of movement of a human body), recognized by a classification model trained using labeled data collected earlier by the subject.”)
Further, WU teaches “a representative pattern data generation unit configured to generate representative pattern data identified as a representative value of the behavior in the refined dataset and store the representative pattern data in a database for representative pattern sample data”:
([Page 851, Col. 2, Paragraph 5] “The proposed CARMUS framework tackles the blackouts issue by exploring the daily repeated pattern of activity data inspired by the daily routine of human activities [15] and the circadian rhythms [16] (representative pattern data). More specifically, we obtain a data matrix by aligning different days’ data with the same time of the day, and estimate the missing values by referring to data collected in previous days in this matrix (the representative pattern data is clearly stored in a database in order to be retrieved, and is being used to refine data here, which was generated based on previously collected data). Through this technique, imputation algorithms designed for multi-dimensional time series [13], [14] can naturally fit into our framework to process data collected in multiple days. Experiment results suggest by increasing the number of days from 1 to 4, the activity recognition accuracy increases from 83.9% to 90.6% on the benchmarking data set, showing the effectiveness of our design.”)
WU fails to explicitly teach “a learning model update unit configured to analyze a similarity of the refined dataset and, based on a result of the analysis, perform learning to generate the behavior recognition model; and an information output unit configured to express a corrected behavior recognition result to a user.”, “a refined data set generation unit configured to divide the sensor data corresponding to a corrected recognition result (a label) to generate the refined dataset including a pair of [label, sensor data].”, “wherein the learning model update unit includes: a dataset similarity analysis unit configured to analyze a similarity of a dataset used for learning; and a behavior recognition model generator configured to analyze a similarity of a dataset and, on the basis of the result of the analysis, perform learning to generate the behavior recognition model”, “wherein the behavior recognition model is configured to, through learning and optimization being performed with a new refined dataset, have various parameters (a weight, a bias, etc.), a layer having a learnable parameter (a convolutional layer, a linear layer, etc.), a value of a registered buffer, and an optimizer and hyperparameter thereof changed to reflect the new refined dataset.”, & “wherein the learning model update unit performs the learning only when a similarity between a dataset having previously been used and the refined dataset is less than or equal to a threshold.”
However, analogous art, KR450 does teach “a learning model update unit configured to analyze a similarity of the refined dataset and, based on a result of the analysis, perform learning to generate the behavior recognition model”:
([Description-of-Embodiments, Paragraphs 15-17] “The unit behavior recognition unit 140 can recognize the unit behavior by classifying the extracted feature information into a unitary action by applying it to a predetermined learning algorithm learned as learning data (Here, the data is classified by applying it to the predetermined learning algorithm learned as learning data, that is, comparing the similarity of it with what is learned already). The learning data is generated in advance as data for learning unit behaviors, and the recognized unit behaviors can be updated to the learning data (thus, the newly recognized data can be used for learning, for the model). The predetermined learning algorithm may be any one of Bayesian, Support Vector Machine (SVM), and Decision Tree algorithm. However, these learning algorithms are merely examples, and the present invention is not limited to this, and various other learning algorithms can be applied.
The representative action recognition unit 150 can recognize at least one representative action based on the unit behavior recognized for each unit data 11 and 12. At this time, since the user can continuously perform two or more representative actions for a long time, if the time interval of the collected sensor data 10 is long, the representative actions recognized through the sensor data 10 can be two or more. Also, the representative behavior recognition unit 150 can recognize the representative behavior of the user by applying various techniques.
For example, when the unit behaviors recognized for the eleven unit data 11 and 12 illustrated in FIG.3 are 'catching a snack', 'setting a food', 'putting in a mouth', ' If you say 'chopsticks catch', 'side dishes',' put in your mouth ',' put down chopsticks', 'catch snipers',' It is possible to recognize the representative behavior as 'meals' by classifying the learning data according to the learning algorithm learned through the learned learning data”)
And further:
([Description-of-Embodiments, Paragraphs 21-23] “The updating unit 160 may update the recognized unit behavior or representative behavior, or both the unit behavior and the representative behavior, to the learning data. The learning data may be stored in the database 170. The updating unit 160 may determine in advance whether to update the recognized representative behavior and each unit behavior to the learning data. This is to increase the quality of learning by accurately determining in advance whether the recognized behaviors are worth learning and updating according to the results.
For example, the update unit 160 may request the learning value determination module to determine whether the recognized representative action or unit action is the learning value, and may receive the result from the learning value determination module to determine whether to update the learning value. The learning value determination module may be implemented in advance by the user through the preprocessing process and constructed in the representative behavior recognition apparatus 100. [Or may be constructed in an external system separately from the representative action recognition apparatus 100.
As another example, when the representative action or the unit behavior is recognized, the update unit 160 may inquire whether the user is learning or not, such as 'Would you like to learn this content?' have. If the unit behavior (e.g., a sub-behavior of a meal) is recognized, the update unit 160 provides the unit behavior to the user, receives an accurate representative action (e.g., meal) To be updated. In addition, it is also possible to provide a unit action and a representative action to a user at the same time to inquire whether the representative action of the unit action is correct.”)
Further, KR450 teaches “an information output unit configured to express a corrected behavior recognition result to a user”:
([Description-of-Embodiments, Paragraph 23] “As another example, when the representative action or the unit behavior is recognized, the update unit 160 may inquire whether the user is learning or not, such as 'Would you like to learn this content?' have. If the unit behavior (e.g., a sub-behavior of a meal) is recognized, the update unit 160 provides the unit behavior to the user, receives an accurate representative action (e.g., meal) To be updated. In addition, it is also possible to provide a unit action and a representative action to a user at the same time to inquire whether the representative action of the unit action is correct.”) Here, the unit behavior is expressed to the user as a corrected recognition result.
Further, KR450 teaches “wherein the learning model update unit includes: a dataset similarity analysis unit configured to analyze a similarity of a dataset used for learning; and a behavior recognition model generator configured to analyze a similarity of a dataset and, on the basis of the result of the analysis, perform learning to generate the behavior recognition model”
([Description-of-Embodiments, Paragraphs 15-17] “The unit behavior recognition unit 140 can recognize the unit behavior by classifying the extracted feature information into a unitary action by applying it to a predetermined learning algorithm learned as learning data (Here, the data is classified by applying it to the predetermined learning algorithm learned as learning data, that is, comparing the similarity of it with what is learned already). The learning data is generated in advance as data for learning unit behaviors, and the recognized unit behaviors can be updated to the learning data (thus, the newly recognized data can be used for learning, for the model). The predetermined learning algorithm may be any one of Bayesian, Support Vector Machine (SVM), and Decision Tree algorithm. However, these learning algorithms are merely examples, and the present invention is not limited to this, and various other learning algorithms can be applied.
The representative action recognition unit 150 can recognize at least one representative action based on the unit behavior recognized for each unit data 11 and 12. [ At this time, since the user can continuously perform two or more representative actions for a long time, if the time interval of the collected sensor data 10 is long, the representative actions recognized through the sensor data 10 can be two or more. Also, the representative behavior recognition unit 150 can recognize the representative behavior of the user by applying various techniques.
For example, when the unit behaviors recognized for the eleven unit data 11 and 12 illustrated in FIG.3 are 'catching a snack', 'setting a food', 'putting in a mouth', ' If you say 'chopsticks catch', 'side dishes',' put in your mouth ',' put down chopsticks', 'catch snipers',' It is possible to recognize the representative behavior as 'meals' by classifying the learning data according to the learning algorithm learned through the learned learning data”)
And further:
([Description-of-Embodiments, Paragraphs 21-23] “The updating unit 160 may update the recognized unit behavior or representative behavior, or both the unit behavior and the representative behavior, to the learning data. The learning data may be stored in the database 170. The updating unit 160 may determine in advance whether to update the recognized representative behavior and each unit behavior to the learning data. This is to increase the quality of learning by accurately determining in advance whether the recognized behaviors are worth learning and updating according to the results.
For example, the update unit 160 may request the learning value determination module to determine whether the recognized representative action or unit action is the learning value, and may receive the result from the learning value determination module to determine whether to update the learning value. The learning value determination module may be implemented in advance by the user through the preprocessing process and constructed in the representative behavior recognition apparatus 100. [Or may be constructed in an external system separately from the representative action recognition apparatus 100.
As another example, when the representative action or the unit behavior is recognized, the update unit 160 may inquire whether the user is learning or not, such as 'Would you like to learn this content?' have. If the unit behavior (e.g., a sub-behavior of a meal) is recognized, the update unit 160 provides the unit behavior to the user, receives an accurate representative action (e.g., meal) To be updated. In addition, it is also possible to provide a unit action and a representative action to a user at the same time to inquire whether the representative action of the unit action is correct.”)
It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of WU with the teachings of KR450 because both references teach of methods for using, updating, and training behavior recognition models.
One of ordinary skill in the art would be motivated to do so because, as KR450 points out in its background, “Recently, many technologies for recognizing user 's behavior through various sensor devices have appeared. However, as a limitation of recognition technology, simple purpose products such as the measurement of the momentum in terms of health are mainly released. The recognition of the user's motion through the sensor is mostly restricted to the amount of exercise such as walking or running, which is also less accurate.” Incorporating methods to dynamically learn to detect more complex behaviors greatly expands use-cases.
WU in view of KR450 still fails to explicitly teach “a refined data set generation unit configured to divide the sensor data corresponding to a corrected recognition result (a label) to generate the refined dataset including a pair of [label, sensor data].”, “wherein the behavior recognition model is configured to, through learning and optimization being performed with a new refined dataset, have various parameters (a weight, a bias, etc.), a layer having a learnable parameter (a convolutional layer, a linear layer, etc.), a value of a registered buffer, and an optimizer and hyperparameter thereof changed to reflect the new refined dataset.” & “wherein the learning model update unit performs the learning only when a similarity between a dataset having previously been used and the refined dataset is less than or equal to a threshold.” However, analogous art, MOJARAD, does teach “a refined data set generation unit configured to divide the sensor data corresponding to a corrected recognition result (a label) to generate the refined dataset including a pair of [label, sensor data].”:
([Abstract] “One of the main objectives of Ambient Assisted Living (AAL) systems is to proactively provide intelligent services to improve the quality of people's lives in terms of autonomy, safety, and well-being. Designing AAL systems that can autonomously monitor human's activities and provide assistance services poses several challenges of which Human Activity Recognition (HAR) which is critically important to adapt the assistance services to the user. In this letter, a robust multi-label HAR framework is proposed. The proposed framework is composed of two main modules: (i) activity classification module (initial classification) and (ii) classification error detection and correction module (classification correction to create refined data). In the first module, machine-learning models are used to predict human activities. Since these models may produce predictions with errors, there is a requirement to detect and correct these errors. The classification error detection and correction module is based on two acyclic directed graphical models and operates in two phases: (i) classification error detection and (ii) classification error correction. The proposed framework is evaluated on the Opportunity dataset, a benchmark and a unique dataset for multi-label human daily living activity recognition. The obtained results demonstrate the ability of the proposed framework to improve the performances of HAR.”)
And further:
([Table II] “Here, NB is the base model’s classifications, while the “Proposed Framework shows the refined dataset generated including a label, sensor data result, with much better accuracy and precision.””)
It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of WU in view of KR450 with the teachings of MOJORAD because both references teach of improving the performance and accuracy of behavior recognition models.
One of ordinary skill in the art would be motivated to do so because, as MOJORAD points out at the end of its abstract, “The obtained results demonstrate the ability of the proposed framework to improve the performances of HAR (human activity recognition).”
WU in view of KR450 & MOJORAD fails to explicitly teach “wherein the behavior recognition model is configured to, through learning and optimization being performed with a new refined dataset, have various parameters (a weight, a bias, etc.), a layer having a learnable parameter (a convolutional layer, a linear layer, etc.), a value of a registered buffer, and an optimizer and hyperparameter thereof changed to reflect the new refined dataset.” & “wherein the learning model update unit performs the learning only when a similarity between a dataset having previously been used and the refined dataset is less than or equal to a threshold.” However, analogous art, THORNTON, does teach “wherein the behavior recognition model is configured to, through learning and optimization being performed with a new refined dataset, have various parameters (a weight, a bias, etc.), a layer having a learnable parameter (a convolutional layer, a linear layer, etc.), a value of a registered buffer, and an optimizer and hyperparameter thereof changed to reflect the new refined dataset.”:
([Abstract] “Many different machine learning algorithms exist; taking into account each algorithm’s hyperparameters, there is a staggeringly large number of possible alternatives overall. We consider the problem of simultaneously selecting a learning algorithm (behavior recognition model) and setting its hyperparameters, going beyond previous work that attacks these issues separately. We show that this problem can be addressed by a fully automated approach, leveraging recent innovations in Bayesian optimization. Specifically, we consider a wide range of feature selection techniques (optimizers) (combining 3 search and 8 evaluator methods) and all classification approaches implemented in WEKA’s standard distribution, spanning 2 ensemble methods, 10 meta-methods, 27 base classifiers, and hyperparameter settings for each classifier. On each of 21 popular datasets (various datasets, which result in changed parameters/hyperparameters) from the UCI repository, the KDD Cup 09, variants of the MNIST dataset and CIFAR-10, we show classification performance often much better than using standard selection and hyperparameter optimization methods. We hope that our approach will help non-expert users to more effectively identify machine learning algorithms and hyperparameter settings appropriate to their applications, and hence to achieve improved performance.”)
And further:
([Page 851-852, 5. EVALUATING AUTO-WEKA] “We now describe an experimental study of the performance that can be achieved by Auto-WEKA on various datasets. After specifying our experiment environment, we demonstrate the importance of addressing the algorithm selection and CASH problems, and establish baselines for them (Section 5.2). We evaluate Auto-WEKA’s ability to search its enormous hyperparameter space effectively to find algorithms and hyperparameters with low cross-validation error (finding various optimized hyperparameters based on the dataset) (Section 5.3). Then, we analyze its test performance and address concerns regarding overfitting (Section 5.4). Finally, we provide a synopsis of the classifiers and feature search/evaluators Auto-WEKA chose in our experiments (Section 5.5).
5.1 Experimental setup
We evaluated Auto-WEKA on 21 prominent benchmark datasets (see Table 3): 15 sets from the UCI repository [12]; the ‘convex’, ‘MNIST basic’ and ‘rotated MNIST with background images’ tasks used in [5]; the appentency task from the KDD Cup ’09; and two versions of the CIFAR-10 image classification task [20] (CIFAR-10-Small is a subset of CIFAR-10, where only the first 10 000 training data points are used rather than the full 50 000.) For datasets with a predefined training/test split, we used that split. (optimized based on dataset) Otherwise, we randomly split the dataset into 70% training and 30% test data. We withheld the test data from all optimization method; it was only used once in an offline analysis stage to evaluate the models found by the various optimization methods. We denote datasets with at least 10 000 training data points as ‘large’ and all others as ‘small’.
All of our experiments were run on Linux machines with Intel Xeon X5650 six-core processors, running at 2.66GHz. We enforced a RAM limit of 3GB for classification (3GB of RAM qualifies as a registered buffer); if training a classifier ever exceeded this memory limit, the classifier job was terminated, returning a misclassification rate of 100%. An additional 1GB of RAM was allocated for the SMBO method. We chose these limits to be reasonably close to the resource limitations faced by a typical user of machine learning algorithms. We also limited the training time for each evaluation of a learning algorithm on each fold, to ensure that the optimization method had a chance to explore the search space. Once this training budget for a fold is consumed, Auto-WEKA sends an interrupt to the learning algorithm to terminate as soon as possible, and the (partially) trained model is then evaluated on the validation set to determine the error estimate of the fold. This timeout was set to 150 minutes for classification and 15 minutes for feature search and evaluation in our experiments.4 For each dataset, we ran Auto-WEKA with each hyperparameter optimization algorithm with a total time budget of 30 hours. For each method, we performed 25 runs of this process with different random seeds and then—in order to simulate parallelization on a typical workstation—used bootstrap sampling to repeatedly select 4 random runs and report the performance of the one with best cross-validation performance.
In early experiments, we observed a few cases in which Auto-WEKA’s SMBO method picked hyperparameters that had excellent training performance, but turned out to generalize poorly. To enable Auto-WEKA to detect such overfitting, we partitioned its training set into two subsets: 70% for use inside the SMBO method, and 30% of validation data that we only used after the SMBO method finished.”)
And further: [Table 4]
Here, we can clearly see the various hyperparameters/parameters being changed and adjusted differently based on the various datasets. Further, many of these parameters are values, which are processed by RAM, which qualifies as a value of a registered buffer, which changes.
And further:
([2. Preliminaries] “This work focuses on classification problems: learning a function
PNG
media_image3.png
19
78
media_image3.png
Greyscale
with finite
PNG
media_image4.png
19
14
media_image4.png
Greyscale
. A learning algorithm A maps a set {d1, . . . , dn} of training data points di = (xi, yi)
PNG
media_image5.png
14
12
media_image5.png
Greyscale
PNG
media_image6.png
16
49
media_image6.png
Greyscale
to such a function, which is often expressed via a vector of model parameters. Most learning algorithms A further expose hyperparameters
PNG
media_image7.png
16
47
media_image7.png
Greyscale
, which change the way the learning algorithm
PNG
media_image8.png
18
22
media_image8.png
Greyscale
itself works. For example, hyperparameters are used to describe a description length penalty, the number of neurons in a hidden layer (a learnable layer), the number of data points that a leaf in a decision tree must contain to be eligible for splitting, etc. These hyperparameters are typically optimized in an “outer loop” that evaluates the performance of each hyperparameter configuration using cross-validation.”)
It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of WU in view of KR450 & MOJORAD with the teachings of THORNTON because WU in view of KR450 & MOJORAD teaches refining the dataset of a behavior recognition model while THORNTON teaches various algorithms for updating hyperparameters, parameters, a learnable layer, etc. of the model based on the dataset.
One of ordinary skill in the art would be motivated to do so because, as THORNTON points out in its conclusion, “In this work, we have shown that the daunting problem of combined algorithm selection and hyperparameter optimization (CASH) can be solved by a practical, fully automated tool.”
Further, WU in view of KR450, MOJORAD, & THORNTON fails to explicitly teach “wherein the learning model update unit performs the learning only when a similarity between a dataset having previously been used and the refined dataset is less than or equal to a threshold.” However, analogous art, BIFET, does teach this:
([Page 443, Introduction, Paragraph 1] “Dealing with data whose nature changes over time is one of the core problems in data mining and machine learning. To mine or learn such data (perform learning), one needs strategies for the following three tasks, at least: 1) detecting when change occurs 2) deciding which examples to keep and which ones to forget (or, more in general, keeping updated sufficient statistics), and 3) revising the current model(s) when significant change has been detected. This describes a need to perform learning when the change is “significant”. In other words, when the difference is above a threshold, which equivalates to a similarity being less than or equal to a threshold.”)
And further:
([Page 444, 2.2 First Algorithm, Paragraphs 1-3] “ADWIN (The learning model) keeps a sliding window W with the most recently read xi. Let n denote the length of W,
PNG
media_image9.png
19
28
media_image9.png
Greyscale
the (observed) average of the elements in W, and μw the (unknown) average of μt for t
PNG
media_image10.png
14
11
media_image10.png
Greyscale
W. Strictly speaking, these quantities should be indexed by t, but in general t will be clear from the context. Here, the dataset is being analyzed as it continuously changes.
Algorithm ADWIN is presented in Figure 1. The idea is simple: whenever two “large enough” subwindows of W exhibit “distinct enough” averages (whenever the differences are great enough/similarity is low), one can conclude that the corresponding expected values are different, and the older portion of the window is dropped (The old data is dropped in order to learn from the new data, once it is shown to be dissimilar/changed enough).
The value of
PNG
media_image11.png
12
28
media_image11.png
Greyscale
for a partition W0 · W1 of W is computed as follows: Let n0 and n1 be the lengths of W0 and W1 and n be the length of W, so n = n0 + n1. Let
PNG
media_image9.png
19
28
media_image9.png
Greyscale
0 and
PNG
media_image9.png
19
28
media_image9.png
Greyscale
1 be the averages of the values inW0 and W1, and μW0 and μW1 their expected values. To obtain totally rigorous performance guarantees we define:
PNG
media_image12.png
103
420
media_image12.png
Greyscale
Our statistical test for different distributions in W0 and W1 simply checks whether the observed average in both subwindows differs by more than the threshold
PNG
media_image11.png
12
28
media_image11.png
Greyscale
(We see the similarity between datasets explicitly being evaluated by a threshold here, where the difference is more than a threshold, meaning that the similarity is less than or equal to a threshold.). The role of
PNG
media_image13.png
14
11
media_image13.png
Greyscale
’ is to avoid problems with multiple hypothesis testing (since we will be testing n different possibilities for W0 and W1 and we want global error below
PNG
media_image13.png
14
11
media_image13.png
Greyscale
). Later we will provide a more sensitive test based on the normal approximation that, although not 100% rigorous, is perfectly valid in practice.”)
It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of WU in view of KR450, MOJORAD, & THORNTON with the teachings of BIFET because both references explore optimal methods for updating and maintaining machine learning models’ learning.
One of ordinary skill in the art would be motivated to do so because, as BIFET points out in its conclusion, “This delivers the user from having to choose any parameter (for example, window size), a step that most often ends up being guesswork. So, client algorithms can simply assume that ADWIN stores the currently relevant data. We tested on both synthetic and real datasets, showed that ADWIN2 really adapts its behavior to the characteristics of the problem at hand.”
Regarding claim 3, WU in view of KR450, MOJORAD, THORNTON, & BIFET teaches the limitations of claim 1. Further, WU teaches “wherein the sensor includes one or more of an accelerometer sensor, a gyroscope sensor, a geomagnetic sensor, an electrocardiogram sensor, a heart rate sensor, a respiration sensor, a skin temperature sensor, and a skin conductivity sensor”
[Figure 2]
Both an accelerometer and a gyroscope sensor can be seen as used in the “Physical Sensors” section of Figure 2.
Regarding claim 4, WU in view of KR450, MOJORAD, THORNTON, & BIFET teaches the limitations of claim 1. Further, KR450 teaches “wherein the data pre-processing unit is configured to, for use in pre-processing, manage a database for representative pattern sample data including samples of various representative values matching a behavior targeted for recognition among datasets having been used for training the behavior recognition model”:
([Description-of-Embodiments, Paragraphs 15-17] “The unit behavior recognition unit 140 can recognize the unit behavior by classifying the extracted feature information into a unitary action by applying it to a predetermined learning algorithm learned as learning data (the learning data is predetermined, and is how the classification result is obtained, representing the various values matching behaviors that was used to train the behavior recognition model). The learning data is generated in advance as data for learning unit behaviors, and the recognized unit behaviors can be updated to the learning data.”)
And further:
([Description-of-Embodiments, Paragraphs 20-21] “According to a further aspect, the apparatus 100 may further comprise an update unit 160 and a database 170. (a database is present)
The updating unit 160 may update the recognized unit behavior or representative behavior, or both the unit behavior and the representative behavior, to the learning data. The learning data may be stored in the database 170 (the learning data is within the database). The updating unit 160 may determine in advance whether to update the recognized representative behavior and each unit behavior to the learning data. This is to increase the quality of learning by accurately determining in advance whether the recognized behaviors are worth learning and updating according to the results.”)
And further:
([Description-of-Embodiments, Paragraph 24] “Meanwhile, the database 170 may store sensor data collected by the collecting unit 110. In addition, the unit data extracting unit 120 may extract the unit data, A reference characteristic to be extracted from the unit 130, and a learning algorithm to be used in the unit behavior recognition unit 140 or the representative action recognition unit 150 may be stored”)
Regarding claim 5, WU in view of KR450, MOJORAD, THORNTON, & BIFET teaches the limitations of claim 4. Further, KR450 teaches “wherein the data pre-processing unit is configured to, in response to a behavior that is repeatedly observed in training datasets or identified to be eligible to have a meaning in addition to the behavior targeted for recognition, add data corresponding to the behavior to the database for the representative pattern sample data and manage the database”
([Description-of-Embodiments, Paragraphs 15-17] “The unit behavior recognition unit 140 can recognize the unit behavior by classifying the extracted feature information into a unitary action by applying it to a predetermined learning algorithm learned as learning data (Here, the data is classified by applying it to the predetermined learning algorithm learned as learning data, that is, comparing the similarity of it with what is learned already). The learning data is generated in advance as data for learning unit behaviors, and the recognized unit behaviors can be updated to the learning data (thus, the newly recognized data can be used for learning, for the model). The predetermined learning algorithm may be any one of Bayesian, Support Vector Machine (SVM), and Decision Tree algorithm. However, these learning algorithms are merely examples, and the present invention is not limited to this, and various other learning algorithms can be applied.
The representative action recognition unit 150 can recognize at least one representative action based on the unit behavior recognized for each unit data 11 and 12. [ At this time, since the user can continuously perform two or more representative actions for a long time, if the time interval of the collected sensor data 10 is long, the representative actions recognized through the sensor data 10 can be two or more. Also, the representative behavior recognition unit 150 can recognize the representative behavior of the user by applying various techniques.
For example, when the unit behaviors recognized for the eleven unit data 11 and 12 illustrated in FIG.3 are 'catching a snack', 'setting a food', 'putting in a mouth', ' If you say 'chopsticks catch', 'side dishes',' put in your mouth ',' put down chopsticks', 'catch snipers',' It is possible to recognize the representative behavior as 'meals' by classifying the learning data according to the learning algorithm learned through the learned learning data”)
And further:
([Description-of-Embodiments, Paragraphs 21-23] “The updating unit 160 may update the recognized unit behavior or representative behavior, or both the unit behavior and the representative behavior, to the learning data. The learning data may be stored in the database 170. The updating unit 160 may determine in advance whether to update the recognized representative behavior and each unit behavior to the learning data. This is to increase the quality of learning by accurately determining in advance whether the recognized behaviors are worth learning and updating according to the results.
For example, the update unit 160 may request the learning value determination module to determine whether the recognized representative action or unit action is the learning value, and may receive the result from the learning value determination module to determine whether to update the learning value. The learning value determination module may be implemented in advance by the user through the preprocessing process and constructed in the representative behavior recognition apparatus 100. [Or may be constructed in an external system separately from the representative action recognition apparatus 100.
As another example, when the representative action or the unit behavior is recognized, the update unit 160 may inquire whether the user is learning or not, such as 'Would you like to learn this content?' have. If the unit behavior (e.g., a sub-behavior of a meal) is recognized, the update unit 160 provides the unit behavior to the user, receives an accurate representative action (e.g., meal) To be updated. In addition, it is also possible to provide a unit action and a representative action to a user at the same time to inquire whether the representative action of the unit action is correct.”)
Regarding claim 7, WU in view of KR450, MOJORAD, THORNTON, & BIFET teaches the limitations of claim 1. Further, WU teaches “wherein the behavior recognition unit further includes a recognition model synchronization unit configured to train the behavior recognition model using the training data including a behavior label (a correct answer sheet) and a sensor dataset corresponding to the behavior label”:
([Page 851, Col. 1, Paragraph 3] “2) Preliminary Results: To understand the issue’s impact on continuous activity recognition performance, we evaluate the recognition accuracy using Phone D’s data (sensor data) by comparing the recognized activities against the results obtained by Phone G (ground-truth sensor data) which is always kept on the subject’s body. We consider three activities including idle, walking, and biking, recognized by a classification model (recognition model) trained using labeled data collected earlier by the subject (labeled behavior data collected by a sensor).”)
Further, WU teaches “the behavior recognition unit is repeatedly updated using information about the behavior recognition model received from the learning model update unit later so as to be kept synchronized with latest information”
([Page 851, Col. 2, Paragraph 5] “The proposed CARMUS framework tackles the blackouts issue by exploring the daily repeated pattern of activity data inspired by the daily routine of human activities [15] and the circadian rhythms [16] (continuously repeated data). More specifically, we obtain a data matrix by aligning different days’ data with the same time of the day, and estimate the missing values by referring to data collected in previous days in this matrix (continuously/repeatedly updating using information about the model so as to be kept synchronized with latest information). Through this technique, imputation algorithms designed for multi-dimensional time series [13], [14] can naturally fit into our framework to process data collected in multiple days. Experiment results suggest by increasing the number of days from 1 to 4, the activity recognition accuracy increases from 83.9% to 90.6% on the benchmarking data set, showing the effectiveness of our design.”)
Regarding claim 9, WU in view of KR450, MOJORAD, THORNTON, & BIFET teaches the limitations of claim 1. Further, KR450 teaches “wherein the time-series characteristics of the behavior includes one of. a constraint of an order of transition between behaviors (sequentiality of behaviors) and a causal necessity of transition between behaviors; a transition time between behaviors taking into account a reaction time of a human body and a duration of a behavior (continuity of a behavior); and a movement having a chance of repetitively occurring (periodicity of a behavior).”:
([Description-of-Embodiments, Paragraphs 17-18] “For example, when the unit behaviors recognized for the eleven unit data 11 and 12 illustrated in FIG.3 are 'catching a snack', 'setting a food', 'putting in a mouth', ' If you say 'chopsticks catch', 'side dishes',' put in your mouth ',' put down chopsticks', 'catch snipers',' It is possible to recognize the representative behavior as 'meals' by classifying the learning data according to the learning algorithm learned through the learned learning data.
Alternatively, when the unit behaviors are performed according to a sequential procedure (sequentiality of behaviors), the HMM (Hidden Markov Model) and the FSM (Finite State Machine) can be used to assume the order when the representative actions are meaningful. In the case of the representative behavior 'meal', it is not meaningful when the unit behaviors are[n’t] performed sequentially, but if the unit behaviors are performed sequentially as in the above example, they can be recognized as 'meal'.”)
Regarding claim 13, WU in view of KR450, MOJORAD, THORNTON, & BIFET teaches the limitations of claim 1. Further, WU teaches “wherein the learning model update unit repeats a process of synchronizing with the behavior recognition model of the behavior recognition unit until the similarity between the datasets converges”:
([Abstract] “This paper presents the CARMUS framework for continuous activity recognition (continuous behavior recognition of real-time data) with missing values on smartphones. Besides the power and resource constraints discussed in existing work, our framework is proposed to further tackle the critical issue of missing values during data collection. We demonstrate the issue's impact on continuous recognition through a motivating example, and specify two challenges-blackouts and resource constraints-with respect to smartphone-based sensing and processing platforms. To address the challenges, CARMUS provides a novel framework which involves a light-weight admission control unit and a data imputation unit intuited by the daily repeated pattern and temporal smoothness of human activity data (refining the results by reflecting time-series characteristics such as “the daily repeated pattern” or “temporal smoothness of human activity data based off of a “continuous” process, thus showing it is repeatedly done until the similarity converges). Based on extensive experiments conducted on a real-world data set with 37% of the data missing, we show that the CARMUS framework is effective for achieving an 85.5% recognition accuracy by adopting the state-of-the-art imputation algorithms. (improving a performance of behavior recognition)”) This citation shows that the synchronization process cited above is used “continuously” to cause the learning to synchronize correctly/converge with the behavior recognition unit.
Regarding claim 14, it comprises similar limitations as claim 1 and is rejected under the same rationale with the following addition: WU teaches “A method of refining data and improving a performance of a behavior recognition model by reflecting time-series characteristics of a behavior”:
([Abstract] “This paper presents the CARMUS framework for continuous activity recognition (behavior recognition of real-time data) with missing values on smartphones. Besides the power and resource constraints discussed in existing work, our framework is proposed to further tackle the critical issue of missing values during data collection. We demonstrate the issue's impact on continuous recognition through a motivating example, and specify two challenges-blackouts and resource constraints-with respect to smartphone-based sensing and processing platforms. To address the challenges, CARMUS provides a novel framework which involves a light-weight admission control unit and a data imputation unit intuited by the daily repeated pattern and temporal smoothness of human activity data (refining the results by reflecting time-series characteristics such as “the daily repeated pattern” or “temporal smoothness of human activity data”). Based on extensive experiments conducted on a real-world data set with 37% of the data missing, we show that the CARMUS framework is effective for achieving an 85.5% recognition accuracy by adopting the state-of-the-art imputation algorithms. (improving a performance of behavior recognition)”)
And further:
([Introduction, B. Proposed Approach & Contributions, Paragraph 2] “The proposed CARMUS framework tackles the blackouts issue by exploring the daily repeated pattern of activity data inspired by the daily routine of human activities [15] and the circadian rhythms [16]. More specifically, we obtain a data matrix by aligning different days' data with the same time of the day (time series data and characteristics), and estimate the missing values by referring to data collected in previous days in this matrix (refining data by reflecting time-series characteristics of a behavior). Through this technique, imputation algorithms designed for multi-dimensional time series [13], [14] can naturally fit into our framework to process data collected in multiple days.”)
Regarding claim 15, WU in view of KR450, MOJORAD, THORNTON, & BIFET teaches the limitations of claim 14. Further, claim 15 recites similar additional limitations as claim 13, and is rejected under the same rationale.
Regarding claim 16, WU in view of KR450, MOJORAD, THORNTON, & BIFET teaches the limitations of claim 14. Further, BIFET, does teach “wherein the generating of the behavior recognition model includes: in the analyzing of the similarity of the dataset, receiving a newly generated refined dataset as input and analyzing the similarity on the basis of a dataset having previously been used”
([Page 443, Introduction, Paragraph 1] “Dealing with data whose nature changes over time (analysis of changing, refined, data) is one of the core problems in data mining and machine learning. To mine or learn such data (perform learning), one needs strategies for the following three tasks, at least: 1) detecting when change occurs 2) deciding which examples to keep and which ones to forget (or, more in general, keeping updated sufficient statistics), and 3) revising the current model(s) when significant change has been detected. This describes a need to perform learning when the change is “significant” which correlates to analyzing the similarity on the basis of previously used data.”)
Further, BIFET teaches “when the similarity is less than or equal to a threshold, returning the refined dataset received as input so that the refined dataset is used for recognition model training”
([Page 443, Introduction, Paragraph 1] “Dealing with data whose nature changes over time is one of the core problems in data mining and machine learning. To mine or learn such data (perform learning), one needs strategies for the following three tasks, at least: 1) detecting when change occurs 2) deciding which examples to keep and which ones to forget (or, more in general, keeping updated sufficient statistics), and 3) revising the current model(s) when significant change has been detected. This describes a need to perform learning when the change is “significant”. In other words, when the difference is above a threshold, which equivalates to a similarity being less than or equal to a threshold.”)
And further:
([Page 444, 2.2 First Algorithm, Paragraphs 1-3] “ADWIN (The learning model) keeps a sliding window W with the most recently read xi. Let n denote the length of W,
PNG
media_image9.png
19
28
media_image9.png
Greyscale
the (observed) average of the elements in W, and μw the (unknown) average of μt for t
PNG
media_image10.png
14
11
media_image10.png
Greyscale
W. Strictly speaking, these quantities should be indexed by t, but in general t will be clear from the context. Here, the dataset is being analyzed as it continuously changes.
Algorithm ADWIN is presented in Figure 1. The idea is simple: whenever two “large enough” subwindows of W exhibit “distinct enough” averages (whenever the difference is great enough/similarity is low), one can conclude that the corresponding expected values are different, and the older portion of the window is dropped (The old data is dropped in order to learn from the new data, once it is shown to be dissimilar/changed enough).
The value of
PNG
media_image11.png
12
28
media_image11.png
Greyscale
for a partition W0 · W1 of W is computed as follows: Let n0 and n1 be the lengths of W0 and W1 and n be the length of W, so n = n0 + n1. Let
PNG
media_image9.png
19
28
media_image9.png
Greyscale
0 and
PNG
media_image9.png
19
28
media_image9.png
Greyscale
1 be the averages of the values inW0 and W1, and μW0 and μW1 their expected values. To obtain totally rigorous performance guarantees we define:
PNG
media_image12.png
103
420
media_image12.png
Greyscale
Our statistical test for different distributions in W0 and W1 simply checks whether the observed average in both subwindows differs by more than the threshold
PNG
media_image11.png
12
28
media_image11.png
Greyscale
(We see the similarity between datasets explicitly being evaluated by a threshold here, where the difference is more than a threshold, meaning that the similarity is less than or equal to a threshold.). The role of
PNG
media_image13.png
14
11
media_image13.png
Greyscale
’ is to avoid problems with multiple hypothesis testing (since we will be testing n different possibilities for W0 and W1 and we want global error below
PNG
media_image13.png
14
11
media_image13.png
Greyscale
). Later we will provide a more sensitive test based on the normal approximation that, although not 100% rigorous, is perfectly valid in practice.”)
Regarding claim 17, WU in view of KR450, MOJORAD, THORNTON, & BIFET teaches the limitations of claim 16. Further, WU teaches “wherein the generating of the behavior recognition model includes: additionally using training data having not been used for learning to verify a performance of the behavior recognition model”:
([Page 856, Col. 2, Paragraph 1] “2) Experiment Methodology: Fig. 6 illustrates our experiment methodology. Data collected by Phone D are filtered and imputed by the proposed CARMUS framework to generate an effective and complete data stream. Both the data streams generated by CARMUS and collected by Phone G are fed to the activity recognition models trained using 45 hours of labeled data from the same subjects. Three activities are selected to demonstrate the effectiveness of the system, which include: idle, walking, and biking. Finally, the system’s performance is evaluated by comparing the predicted and ground-truth activity label sequences. Hours without ground-truth data are omitted during the evaluation to make the results comprehensible. Additionally, we study the performance of the system under different configurations with respect to the baseline accuracy of 86.1% which is obtained by only using the effective data and their corresponding ground-truth for evaluation (as did in Sec. I-A). We evaluate the system’s performance by comparing to the baseline accuracy instead of 100% accuracy to eliminate the impact of differences in data collected by the two phones.”)
Further, WU teaches “obtaining high performance even for new data not shown in learning”
([Page 858, VI. Conclusion, Paragraph 2] “Extensive experiments are conducted using a real-world data set with 37% of the data missing. The results suggest the framework, by adopting the DynaMMo [13] imputation algorithm, is effective by maintaining the system’s recognition accuracy to be 85.5% (high performance), comparable to the baseline which only uses the effective data for recognition.”)
Further, THORNTON teaches “tuning various parameters having been used in the model to prevent overfitting of behavior recognition model”
([Page 853, 5.4 Results for Test Performance, Paragraphs 1-2] “The results just shown demonstrate that Auto-WEKA is effective at optimizing its given objective function; however, this is not sufficient to allow us to conclude that it fits models that generalize well. As the number of hyperparameters of a machine learning algorithm grows, so does its potential for overfitting. The use of cross-validation substantially increases Auto-WEKA’s robustness against overfitting, but since its hyperparameter space is much larger than that of standard classification algorithms, it is important to carefully study whether (and to what extent) overfitting poses a problem.
To evaluate generalization, we determined a combination of algorithm and hyperparameter settings
PNG
media_image14.png
16
21
media_image14.png
Greyscale
by running Auto-WEKA as before (cross-validating on the training set), trained
PNG
media_image14.png
16
21
media_image14.png
Greyscale
on the entire training set, and then evaluated the resulting model on the test set. The right portion of Table 4 reports the test performance obtained with all methods.”) Here we see that cross-validation and a combination of algorithm and hyperparameter settings (tuning various parameters) is used to increase “robustness against overfitting”/prevent overfitting.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over WU in view of KR450, MOJORAD, THORNTON, & BIFET as applied to claims above, and further in view of Ferrari, A. et al. “On the Personalization of Classification Models for Human Activity Recognition.” Available at https://ieeexplore.ieee.org/iel7/6287639/8948470/08995531.pdf on February 24 2020 (hereafter, FERRARI)
Regarding claim 6, WU in view of KR450, MOJORAD, THORNTON, & BIFET teaches the limitations of claim 4. WU in view of KR450, MOJORAD, THORNTON, & BIFET fails to explicitly teach “wherein the data pre-processing unit is configured to: when data having been used for the learning includes various users or sensors, manage generalized or general-purpose representative pattern data in a database; and when data having been used for the learning includes data from a specific user or specific sensor, manage personalized or dedicated representative pattern data in a database.” However, analogous art, FERRARI, does teach this:
([Page 32068, B. Personalization in Har, Paragraphs 4-10] “Thus, as users of mobile sensing applications increase in size, the differences between people cause the accuracy of classification to degrade quickly [17].
To face this problem, activity classification models should be able to generalize as much as possible with respect to the final user and the real execution context.
In order to achieve generalizable activity recognition models, three approaches are mainly adopted in literature: subject-independent (generalized or general-purpose), subject-dependent (personalized or dedicated), and hybrid.
The subject-independent (also called impersonal) model (a machine learning model with its own database) does not use the end user data for the development of the activity recognition model. It is based on the definition of a single activity recognition model that must be flexible enough to be able to generalize the diversity between users and it should be able to have good performance once a new user is to be classified. Meaning that this model is trained using data learned from various users or sensors to manage generalized or general-purpose representative pattern data in a database (the model’s own data).
The subject-dependent (also called personal) model (a machine learning model with its own database) only uses the end user data for the development of the activity recognition model. The specific model, being built with the data of the final user, is able to capture her/his peculiarities, thus it should well generalize in the real context. The flaw is that it must be implemented for each end user [39]. (Meaning that this model is trained using data from a specific user or specific sensor to manage personalized or dedicated representative pattern data in a database (the model’s own data))
The hybrid model uses the end user data and the data of the other users for the development of the activity recognition model. In other words the classification model is trained both on the data of the users and on a part of the data of the final user. The idea is that the classifier should recognize easier the activity performed by the final user. This would also qualify as a generalized general-purpose model.
Figure 1 shows a graphical depiction of the three models to better clarify their differences.”)
Figure 1 has been inserted her by the examiner for reference and convenience.
It would be obvious to one of ordinary skill in the art, prior to the effective filing date of the claimed invention, to combine the base reference of WU in view of KR450, MOJORAD, THORNTON, & BIFET with the teachings of FERRARI because both references teach related concepts for the management and training of behavior recognition models.
One of ordinary skill in the art would be motivated to do so because as FERRARI points out in the abstract, “The experiments show that the employment of personalization models improves, on average, the accuracy”.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW LEE LEWIS whose telephone number is (571)272-1906. The examiner can normally be reached Monday: 12:00PM - 4:00PM and Tuesday - Friday: 12:00PM - 9PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tamara Kyle can be reached at (571)272-4241. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Matthew Lee Lewis/ Examiner, Art Unit 2144
/TAMARA T KYLE/ Supervisory Patent Examiner, Art Unit 2144