DETAILED ACTION
This action is responsive to Applicant’s reply filed June 29th 2026. This action is made non-final.
Status of the Claims
Claims 1 and 13 are amended.
Claim 9 is canceled.
Claim 21 is added.
Claim status is currently pending and under examination for Claims 1-8 and 10-21 of which independent claims are 1 and 13.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on June 29, 2026 has been entered.
Response to Amendment
Applicant’s arguments regarding the art rejections are moot in view of the new grounds of rejection necessitated by Applicant’s amendment.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The following are the references relied upon in the rejections below:
Patil (US 20190347327 A1)
Pujari, Subhash Chandra, et al. "Evaluating neural multi-field document representations for patent classification." BIR 2022-Bibliometric-enhanced Information Retrieval: Proceedings of the 12th International Workshop on Bibliometric-enhanced Information Retrieval co-located with 44th European Conference on Information Retrieval (ECIR 2022), April 10th 2022, Stravanger, Norway. CEUR-WS, 2022.
Burns (US 20200226321 A1)
Gopalakrishnan (US 20230134546 A1)
Claims 1-5, 13-17 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Patil / Pujari / Burns / Gopalakrishnan.
Regarding Claim 1, Patil teaches:
A computer-implemented method, comprising ([0002] “methods for automatically assigning one or more labels or tags related to various discussion forum topics”):
receiving, by a data processing system, event data … ([0023] “an inventive computer-implemented system (hereinafter “system” or “present system”) that involves automatically assigning one or more labels (tags), in a hierarchical structure, to discussion topics seen in d2web forums. … the present system crawls various d2web sites and extracts important information from html pages, storing it in a database. … The important information is then parsed from these sites, such as discussion topics, user related information, discussion posts, etc., and the information stored on a database as well as on an elastic search data store.”
A present system (‘data processing system’) collects (‘receives’) and stores html pages (‘event data’).),
wherein the event data comprises a data structure having a plurality of fields and a plurality of indicators delineating each field of the plurality of fields from other fields of the plurality of fields ([0025] “From the collected html pages or documents of the dataset, the present system parses predetermined important fields from the data such as topic title, posts, user name, title posted date, user ratings, number of replies, etc.”
Html pages (‘event data’) have predetermined important fields (‘plurality of fields’) and their corresponding data that make up a data structure of the html pages. Html pages are parsed to identify predetermined important fields, therefore the fields also act as indicators delineating (separating) each field from other important fields (since the fields indicate where data for each field begins and ends).);
processing the event data, including: parsing, by a parser of the data processing system based on the plurality of indicators delineating each field of the plurality of fields, the event data to identify the data structure of the event data ([0065] “Each site has a parser program that extracts important information, such as fields including a topic title, topic content, post content, topic author, post author, author reputation, posted date for each title and posts, etc., from HTML pages, all of which is stored on a database.”
The present system (‘data processing system’) uses parser programs to parse (‘process’) the html pages (‘event data’) based on predetermined important fields (‘plurality of indicators delineating each field’) to identify important fields. By identifying important fields, a data structure of an html page is identified since each field and its corresponding data is identified in the page by parsing.);
and identifying a given field and data content of that given field, by: identifying the plurality of fields from the data structure of the event data including data content corresponding to each field of the plurality of fields (By parsing html pages, a parse program is able to identify from the structure (collection of important fields and corresponding data) of an html page each important field. Each important field (topic title, post content) contains its corresponding data, therefore identifying a given field and data content of that given field.);
inputting, to a machine learning engine, data content of the given field of the plurality of fields ([0027] “Textual feature extraction was performed on topic title using a Doc2vec vectorization technique. Doc2vec is an extension of Word2vec… Word2vec is computationally-efficient predictive model that uses shallow two layer neural network for learning word embeddings.”
A topic title (data content of a topic title important field) is input into a Doc2vec predictive model (‘machine learning engine’).);
generating, by the machine learning engine and from contents of the one or more fields, one or more feature vectors by: identifying, in the data content of the given field, one or more values, each value corresponding to a feature of one or more features ([0023] “Data preprocessing and feature extraction is then performed over the forum discussion topics. … The present system then extracts features from every topic on the forum using Doc2vec.”
Extracted features (‘feature vector’) are generated from a discussion topic (topic title) by using Doc2vec. By extracting features from a topic title, each feature has a corresponding value identified from data content of a topic title field (‘given field’).);
accessing, from the hardware storage device, a plurality of parent indicator candidates and a set of child indicator candidates for each parent indicator candidate ([0033] “In the ground truth, each topic is associated with multiple tags. It was observed that there is a natural hierarchical structure in the set of tags. There are group tags (parent tags) that can be seen as the “broader term” for its set of specific tags (child tags). Nesting specific tags under group tags creates a hierarchy of tags. FIGS. 5A-5C show an example of the hierarchical structure in the training dataset.”
[0043] “an automation script was developed that establishes the hierarchy in topic titles in the ground truth. Now, with hierarchy established in the tags, the classifier provided detailed tag prediction for all the documents.”
A classifier uses hierarchy established in topic titles to generate tag predictions for documents. A training dataset is comprised of ground truth that establishes a hierarchy of parent tags (‘parent indicator candidates’) and child tags (‘child indicator candidates’). To use a hierarchy of tags to generate predictions, the ground truth must be accessed and stored somewhere in memory, therefore accessing, from a hardware storage device, parent indicator candidates (parent tags) and child indicator candidates (child tags) is implied.);
determining, by a parent prediction model of the machine learning engine, a parent probability value for each parent indicator candidate of the plurality of parent indicator candidates based on the one or more feature vectors ([0023] “The present system then extracts features from every topic on the forum using Doc2vec. Finally, in the fourth module, different machine classifiers are used to assign multiple tags to each topic, so that their performances can be compared.”
[0033] “In the ground truth, each topic is associated with multiple tags. It was observed that there is a natural hierarchical structure in the set of tags. There are group tags (parent tags) that can be seen as the “broader term” for its set of specific tags (child tags). Nesting specific tags under group tags creates a hierarchy of tags.”
Patil discloses Algorithm 1 in [0052] (reproduced below) describing an algorithm for adding parent tags and removing child tags. Extracted features from each topic (‘feature vector’) are used to assign (predict) tags to each topic by using a machine classifier (‘parent prediction model’). For each tag in a prediction list of tags (‘parent indicator candidates’), a prediction probability p(t) (‘parent probability value’) is determined by the machine classifier and compared to parent threshold value α (Lines 6-8).
PNG
media_image1.png
745
1205
media_image1.png
Greyscale
);
determining, by a child prediction model of the machine learning engine, a child probability value for each child indicator candidate of the set of child indicator candidates ([0044] “the classifier should predict the hierarchical representation of tags, for example, for a predicted child tag every parent tag should also be predicted”
A machine classifier is used to predict both parent and child labels, therefore the machine classifier is both a parent prediction model and a child prediction model. In Algorithm 1, for each tag in a prediction list of tags (‘child indicator candidates’), a prediction probability p(t) (‘child probability value’) is determined by the machine classifier and compared to child threshold value β (Lines 3-5).);
tagging, by the data processing system, the event data with: each parent indicator candidate having a parent probability value that satisfies the parent threshold ([0023] “machine classifiers are used to assign multiple tags to each topic”
[0081] “the desire was to regulate the way parent tags were added and child tags removed based on the probability of the predicted tag. The “add parent threshold” (α) or the “remove child threshold values” (β) are used in order to decide on whether to add or remove a particular tag, respectively.”
In Algorithm 1, for each tag in a prediction list of tags (‘parent indicator candidates’), a prediction probability p(t) (‘parent probability value’) is compared to parent threshold value α (Lines 6-8). If the prediction probability p(t) for a tag is greater than the parent threshold value (satisfying the parent threshold), then the tag is added to an updated prediction list. The updated prediction list is then returned (last line) and the tags in the updated prediction list are assigned to a topic (‘event data’).);
and for each tagged parent indicator candidate, each child indicator candidate of the set of child indicator candidates having a child probability that satisfies a child threshold ([0081] “the desire was to regulate the way parent tags were added and child tags removed based on the probability of the predicted tag. The “add parent threshold” (α) or the “remove child threshold values” (β) are used in order to decide on whether to add or remove a particular tag, respectively.”
In Algorithm 1, for each tag in a prediction list of tags (‘child indicator candidates’), a prediction probability p(t) (‘child probability’) is compared to child threshold value β (Lines 3-5). If the prediction probability p(t) for a child tag is less than the child threshold value, then the child tag is removed from an updated prediction list. Child tags with a probability greater than child threshold value β are kept in the updated prediction list (therefore satisfying a child threshold). The updated prediction list is then returned and the child tags (along with parent tags) in the updated prediction list are assigned to a topic (‘event data’).);
However, Patil does not teach evaluating a corresponding set of child indicator candidates based on a parent indicator candidate having a probability value that satisfies or does not satisfy a parent threshold, which is taught by Pujari:
for each parent indicator candidate having a probability value that satisfies a parent threshold, determining, by a child prediction model of the machine learning engine, a child probability value for each child indicator candidate of the set of child indicator candidates ((P. 20, Sec. 4.3, ¶1) “we use TwistBytes [15], a Local Classifier per Node (LCN) approach which trains a Support Vector Machine (SVM) model [29] as a base classifier for each label in the class hierarchy. … The tree is traversed from the root node to leaf nodes during prediction, predicting a label at each hop using the label-specific classifier. A child label classifier is traversed if the probability of the parent label is more than a user-defined threshold.”
A tree is used to organize parent and child labels in a hierarchy. At each node of the tree, a classifier is trained to predict a label for the node. A parent node uses a trained SVM model to generate a probability of a parent label (‘parent indicator candidate’). If the probability of the parent label is more than a user-defined threshold (‘parent threshold’ is satisfied), then a child label classifier is traversed. When a child label classifier is traversed, a classifier (‘child prediction model’) of the child node can generate a prediction (child probability) for a child label (‘set of child indicator candidates’).);
for each parent indicator candidate having a probability value that does not satisfy a parent threshold, forgoing evaluating the corresponding set of child indicator candidates (A parent node uses a trained SVM model to generate a probability of a parent label (‘parent indicator candidate’). When the probability of the parent label is less than the user-defined threshold (‘parent threshold’), the threshold is not satisfied and the child label classifier is not traversed. Since a child label classifier is not traversed, a classifier of the child node cannot predict a child label (‘set of child indicator candidates’), therefore forgoing evaluating a corresponding set of child indicator candidates.);
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the tagging method of Patil with the threshold disclosed by Pujari to compare a probability of a parent label to a threshold. By comparing a probability of a parent label to a threshold, prediction of child labels for parent labels that fail to satisfy a probability threshold can be skipped, thereby saving time and computational resources.
However, the combined tagging method of Patil / Pujari does not teach event data representing a medical safety event and updating a machine learning engine based on tagged, processed event data stored in a hardware storage device, which is taught by Burns:
receiving, by a data processing system, event data representing a medical safety event ([0021] “text processor 12 is configured to receive an input record, where the input record represents a medical procedure performed on a patient and the input record includes a text description describing the medical procedure”);
storing the tagged event data in the hardware storage device ([0023] “the input record with an assigned billing code is passed directly to a billing system 17 for processing. In other embodiments, the input record with the assigned billing code are reviewed manually by a billing specialist on a user interface of a computing device 18 before being passed on the billing system 17. The billing specialist may elect to confirm the assignment made by the system or change the assignment made the system”);
and updating the machine learning engine based on the tagged, processed event data stored in the hardware storage device ([0024] “the assignment of the billing code to an input record (either by the system or manually) is used as feedback to improve the machine learning models. That is, the models can be re-trained and/or updated based on the input records with assigned billing codes. Additionally or alternatively, the models can be re-trained and/or updated using feedback from a billing specialist. During a validation process, the billing specialist can indicate whether a billing code assignment was accurate or not and, if not, provide a reason. The feedback from the billing specialist in turn is represented as a vector that is used to re-train the machine leaning models”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined tagging method of Patil / Pujari with the training technique disclosed by Burns to use tagged medical event data to retrain a classification model. By using tagged medical event data to retrain a classification model, tagged medical event data can be corrected and used to retrain the model, thereby allowing the model to learn from corrected tags and generate accurate tags to keep medical data organized over time.
Furthermore, the combined tagging method of Patil / Pujari / Burns does not teach calculating, based on a database of historical event data stored in a hardware storage device, a frequency score corresponding to each feature, which is taught by Gopalakrishnan:
and calculating, based on a database of historical event data stored in a hardware storage device, a frequency score corresponding to each feature of the one or more features ([0057] “ML service 112 includes an NLP engine that converts text-based features into numerical values. For instance, a numerical value for a token may be a score that represents what the word means to the log entry versus what the word means to a list of historical events. An example approach for assigning a score is to compute a term frequency inverse-document frequency (TF-IDF) score. With TF-IDF, the score for a token increases proportionally to the frequency the token appears in a log record offset by the number of logs that include the token”
[0130] “the microservice application may generate and provide an output based on input that identifies, locates, or provides historical data”
A microservice application receives and provides historical data, therefore it is implied historical data is stored in an accessible database (‘hardware storage device’).),
the frequency score indicating a frequency at which the feature appears in the historical data ([0061] “The process further determines a frequency of the textual token in a list of historical events (operation 206). For example, the process may determine a frequency of the token in a list of log records in the past five days or over some other timeframe”);
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined tagging method of Patil / Pujari / Burns with the TF-IDF scores disclosed by Gopalakrishnan to calculate a TF-IDF score for each word feature based on historic data. By calculating a TF-IDF score for each word feature based on historic data, the frequency of a word across several historical documents can be used to determine a word’s importance, thereby filtering out words that appear frequently and elevating the importance of rare words.
Regarding Claims 2 and 14, the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan teaches:
the computer-implemented method of claim 1, wherein the one or more feature vectors comprise the one or more features and one or more corresponding scores including the frequency score for each feature of the one or more features (Gopalakrishnan discloses a TF-IDF score (‘frequency score’) is computed for each token (‘feature’), “ML service 112 includes an NLP engine that converts text-based features into numerical values. For instance, a numerical value for a token may be a score that represents what the word means to the log entry versus what the word means to a list of historical events. An example approach for assigning a score is to compute a term frequency inverse-document frequency (TF-IDF) score. With TF-IDF, the score for a token increases proportionally to the frequency the token appears in a log record offset by the number of logs that include the token” [0057].
[0029] “The model may use NLP to convert text associated with account activity to numerical vectors, where the vectors include scores and/or other numerical values computed based on the meaning of the converted text. The model may further include a set of classifiers trained to learn patterns in the numerical vectors that are predictive of a network attack. The model may assign labels to events based on the predicted likelihood that the event is an attack”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan with the TF-IDF scores disclosed by Gopalakrishnan to use TF-IDF scores to tag event data. By using TF-IDF scores to tag event data, TF-IDF can be used to identify rare words and train a classifier to associate rare words with a label, thereby ignoring common words and increasing tag prediction accuracy.
Regarding Claims 3 and 15, the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan teaches:
the computer-implemented method of claim 2, wherein the one or more values comprise a plurality of words that describe the medical safety event (Burns discloses feature vectors are comprised of strings (‘one or more features’), “a feature vector is constructed at 33 by extracting one or more features from the input record. In this example embodiment, elements in the feature vector include each string in the text description” [0042].
Burns discloses strings are comprised of words that describe medical procedures (‘medical safety events’), “to aid in the processing of procedural text and to decrease vocabulary size, the text description from the input record was processed into a standardized form. Because the text description of the medical procedure is typically hand-entered, it is subject to misspellings and frequently contains medical abbreviations and acronyms. First, misspelled words in the text description are corrected” [0030].),
wherein the one or more features are determined from the plurality of words (Burns discloses “each string in the standardized form of the text description is an element in the feature vector” [0026].),
and wherein the frequency score corresponding to each feature of the one or more features comprises a term frequency-inverse document frequency (TF-IDF) score (Gopalakrishnan discloses a TF-IDF score (‘frequency score’) is computed for each word token (‘feature’), “ML service 112 includes an NLP engine that converts text-based features into numerical values. For instance, a numerical value for a token may be a score that represents what the word means to the log entry versus what the word means to a list of historical events. An example approach for assigning a score is to compute a term frequency inverse-document frequency (TF-IDF) score. With TF-IDF, the score for a token increases proportionally to the frequency the token appears in a log record offset by the number of logs that include the token” [0057].).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan with the feature vectors disclosed by Burns to use a machine learning model to tag medical event data. By using a machine learning model to tag medical event data, medical event data can be automatically tagged based on extracted words, thereby automating medical event data labeling and organization.
Regarding Claims 4 and 16, the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan teaches:
The computer-implemented method of claim 2, wherein determining the one or more indicators from the plurality of indicator candidates comprises: determining a probability value corresponding to each indicator candidate based on the one or more features and the one or more corresponding scores (Gopalakrishnan discloses “Based applied classifier, the process generates an output based on the predicted likelihood that the current account activity constitutes an attack … the output includes a label, such as Red, Amber, or Green” [0091]
[0029] “The model may use NLP to convert text associated with account activity to numerical vectors, where the vectors include scores … The model may further include a set of classifiers trained to learn patterns in the numerical vectors that are predictive of a network attack. The model may assign labels to events based on the predicted likelihood that the event is an attack”).
See [0057] describing a TF-IDF score (‘corresponding score’) is found for each word (‘feature’) in an event.
A classifier uses TF-IDF scores and text (words) to determine a predicted likelihood (‘probability value’) to assign account activity a Red, Amber, or Green label, thereby determining a probability value corresponding to each indicator candidate (Red, Amber, or Green labels) based on features and corresponding scores.);
comparing the probability value corresponding to each indicator candidate with a probability threshold value (Patil discloses [0081] “the desire was to regulate the way parent tags were added and child tags removed based on the probability of the predicted tag. The “add parent threshold” (α) or the “remove child threshold values” (β) are used in order to decide on whether to add or remove a particular tag, respectively.”
See in Algorithm 1, for each tag in a prediction list of tags (‘parent indicator candidates’), a prediction probability p(t) (‘parent probability value’) is compared to parent threshold value α (Lines 6-8). For each tag in a prediction list of tags (‘child indicator candidates’), a prediction probability p(t) (‘child probability value’) is compared to child threshold value β (Lines 3-5).);
and selecting the one or more indicators whose corresponding probability values are greater than the probability threshold value (In Algorithm 1, if the prediction probability p(t) for a parent tag is greater than the parent threshold value, then the parent tag is added (‘selected’) to an updated prediction list. If the prediction probability p(t) for a child tag is less than the child threshold value, then the child tag is removed from an updated prediction list. Child tags with a probability greater than child threshold value β are kept in the updated prediction list (therefore selecting child labels with probability values greater than a child threshold).).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan with the TF-IDF scores disclosed by Gopalakrishnan to use TF-IDF scores to tag event data. By using TF-IDF scores to tag event data, TF-IDF can be used to identify rare words and train a classifier to associate rare words with a label, thereby ignoring common words and increasing tag prediction accuracy.
Regarding Claims 5 and 17, the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan teaches:
The computer-implemented method of claim 1, further comprising: training the machine learning engine with a set of sample event data ([0026] “The ground truth to build a machine learning model may be hand labeled by field experts. The present system has 2,046 labeled topics as a training set and they belong to 226 unique tags.”),
wherein the sample event data are tagged with one or more sample indicators (The ground truth in a training set (‘sample event data’) contain labeled topics (‘event data’) and their tags (‘sample indicators’).).
Regarding claim 13, the rejection of claim 1 is incorporated. The difference in scope being
A non-transitory computer-readable medium storing program instructions that cause a data processing system to perform operations comprising ([0089] “a computer program product (e.g., a computer program tangibly or non-transitorily embodied in a machine-readable medium and including instructions for execution by, or to control the operation of, a data processing apparatus, such as, for example, one or more programmable processors or computers).”).
Regarding Claim 21, the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan teaches:
The computer-implemented method of claim 1, further comprising: determining the parent threshold for each parent indicator candidate by: setting a target range for a precision score corresponding to the parent prediction model ([0082] “Removing child tags reduces false positives, increasing the precision score, while adding parent tags reduces false negatives, increasing the recall score. The above experiment was repeated for different “remove child threshold” (β) and “add parent threshold” (α) values. As shown in FIG. 10, with an increase in “add parent threshold” (α) value, the recall score increased and precision score decreased. … The aim was to find an optimum increase in F1 score without drastically affecting the precision score by tuning the (α) and (β) values. The best F1 score was achieved when β=0.5 and α=0.9.”
[0074] “The performance of the classifier models is evaluated based on four metrics—precision, recall, F1 score, and percentage of documents titles with at least one correct prediction tag. … F1 measure is the harmonic mean of precision and recall score. The precision, recall, and F1 scores were calculated for individual tags”
An optimal F1 score is found by tuning parent threshold α. The optimal F1 score is a harmonic mean of precision and recall, therefore recall and precision for a parent label need to be adjusted so that recall and precision can both be high. Since an increase of recall decreases precision, an optimal value of parent threshold α needs to be found that balances both recall and precision, therefore setting a target range for a precision score corresponding to a parent prediction model (machine classifier). The target range for precision being a precision that is high, but low enough not to decrease recall.);
setting a target range for a recall score corresponding to the parent prediction model (Since an increase of recall decreases precision, an optimal value of parent threshold α needs to be found that balances both recall and precision, therefore setting a target range for a recall score corresponding to a parent prediction model. The target range for recall being a recall that is high, but low enough not to decrease precision.);
and determining the parent threshold such that the precision score corresponding to the parent prediction model falls within the target range of precision scores and the recall score corresponding to the parent prediction model falls within the target range of recall scores (An optimal value of parent threshold α (α=0.9) is found by balancing recall and precision to find an optimal F1 score. The optimal F1 score is found when recall is high, but low enough not to decrease precision (recall falls within target range of recall scores) and when precision is high, but low enough not to decrease recall (precision falls within target range of precision scores).);
determining the child threshold for each child indicator candidate by: setting a target range for a precision score corresponding to the respective child prediction model ([0082] “Removing child tags reduces false positives, increasing the precision score, while adding parent tags reduces false negatives, increasing the recall score. The above experiment was repeated for different “remove child threshold” (β) and “add parent threshold” (α) values. … Conversely, with an increase in “remove child threshold” value (β), the precision score increased and recall score decreased. The aim was to find an optimum increase in F1 score without drastically affecting the precision score by tuning the (α) and (β) values. The best F1 score was achieved when β=0.5 and α=0.9.”
An optimal F1 score is found by tuning child threshold β. The optimal F1 score is a harmonic mean of precision and recall, therefore recall and precision for a child label need to be adjusted so that recall and precision can both be high. Since an increase of recall decreases precision (and an increase of precision decreases recall), an optimal value of child threshold β needs to be found that balances both recall and precision, therefore setting a target range for a precision score corresponding to a child prediction model (machine classifier). The target range for precision being a precision that is high, but low enough not to decrease recall.);
setting a target range for a recall score corresponding to the respective child prediction model (Since an increase of recall decreases precision (and an increase of precision decreases recall), an optimal value of child threshold β needs to be found that balances both recall and precision, therefore setting a target range for a recall score corresponding to a child prediction model. The target range for recall being a recall that is high, but low enough not to decrease precision.);
and determining the child threshold such that the precision score corresponding to the respective child prediction model falls within the target range of precision scores and the recall score corresponding to the respective child prediction model falls within the target range of recall scores (An optimal value of child threshold β (β=0.5) is found by balancing recall and precision to find an optimal F1 score. The optimal F1 score is found when recall is high, but low enough not to decrease precision (recall falls within target range of recall scores) and when precision is high, but low enough not to decrease recall (precision falls within target range of precision scores).).
The following are the references relied upon in the rejections below:
Kabeya (US 20190122144 A1)
Claims 6-8 and 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Patil / Pujari / Burns / Gopalakrishnan / Kabeya.
Regarding Claims 6 and 18, the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan teaches:
the computer-implemented method of claim 5, however the combination does not teach obtaining sample feature vectors from the sample event data, obtaining an indicator vector, and inputting the vectors into a logistic regression classifier which is taught by Kabeya:
wherein training the machine learning engine with the set of sample event data comprises: obtaining one or more sample feature vectors from the sample event data ([0041] “event record system 150 may collect event information originating from one or more event sources, and record the collected event information to the event collection database 120 together with its timestamp as a data record. Such event sources may include, but not limited to, electronic medical record systems”
See Figure 4 (reproduced below) depicting obtaining a set s of timestamps (‘sample feature vector’) from data records (‘sample event data’).);
obtaining an indicator vector from the sample event data, wherein the indicator vector comprises a plurality of fields indicating a presence or absence of the plurality of indicator candidates (Kabeya discloses Figure 4 (reproduced below) depicting an input vector (‘indicator vector’) comprised of 1’s and 0’s, with 1 indicating a label (‘indicator candidate’) exists in the input set (‘sample event data’) and 0 indicating a label does not exist in the input set.
PNG
media_image2.png
648
1126
media_image2.png
Greyscale
Kabeya discloses input vector as set U “there is the input vector 220 including a plurality of elements that corresponds to the predetermined label set
L
=
{
L
1
,
.
.
.
,
L
N
}
. Each element
u
n
has a value representing at least whether or not a corresponding label
L
n
is observed at least once actually in the set of the data records 210. The input vector generation module 112 may set the element
u
n
by one
u
n
=1) when the corresponding label
L
n
exists in the set of the data records
{
l
1
,
.
.
.
,
l
N
}
. The input vector generation module 112 may set the element
u
n
by zero (
u
n
=0) when the corresponding label
L
N
does not exist in the set of the data records
{
l
1
,
.
.
.
,
l
N
}
” [0055].);
and inputting the one or more sample feature vectors and the one or more indicator vectors to a logistic regression classifier to obtain a prediction model (Kabeya discloses a regression model receives as input an input vector (‘indicator vector’) and a set of timestamps (‘sample feature vectors’), “the regression model 160 includes an input layer 162 corresponding to the predetermined label set L; and an output layer 168 configured to output the probability for a given target timestamp t*; and a network structure provided therebetween. The input layer 162 is configured to receive the input vector u and the representative timestamps s that are obtained from the set of the data records” [0059-0060]. See Figure 5 depicting the inputs of a regression model.
Kabeya discloses a regression model’s sigmoid function can be used as a classifier (‘prediction model’), “the regression model 160 employed in the described embodiment can be seen as an extension of a binary logistic regression model where a binary dependent variable (the target label exists or does not exist) is used and a sigmoid function is used as the output function, which can be used as classifier” [0067].).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan with the technique disclosed by Kabeya to identify labels that do not exist in a data set. By identifying which labels are present or absent in data, a model can learn to better distinguish between features that indicate the presence of a label and features that indicate the absence of label, thereby yielding more accurate classifications.
Regarding Claims 7 and 19, the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan / Kabeya teaches:
the computer-implemented method of claim 6, wherein training the machine learning engine with the set of sample event data further comprises: obtaining one or more prediction metrics from a set of test event data (Patil discloses [0076] “In all the experiments, a training set with topic titles to perform 10-fold cross validation is used to validate the models. With 10-fold cross validation, the ground truth data is randomly partitioned into 10 equal subsample buckets. Out of 10 buckets, one bucket is used for testing the model and the rest of the ground truth is used for training the model. Each of the 10 bucket subsamples are used for testing the model in cross validation. The aggregate of the accuracy scores of k scores is used as the final accuracy score”
[0033] “In the ground truth, each topic is associated with multiple tags”);
and determining a probability threshold value for the prediction model ([0082] “Removing child tags reduces false positives, increasing the precision score, while adding parent tags reduces false negatives, increasing the recall score. The above experiment was repeated for different “remove child threshold” (β) and “add parent threshold” (α) values. … The aim was to find an optimum increase in F1 score without drastically affecting the precision score by tuning the (α) and (β) values. The best F1 score was achieved when β=0.5 and α=0.9.”
See [0081] describing parent and child threshold values are compared to a probability of each predicted tag.).
Regarding Claims 8 and 20, the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan / Kabeya teaches:
The computer-implemented method of claim 7, wherein the one or more prediction metrics comprise: a precision threshold value and a recall threshold value ([0082] “Removing child tags reduces false positives, increasing the precision score, while adding parent tags reduces false negatives, increasing the recall score. The above experiment was repeated for different “remove child threshold” (β) and “add parent threshold” (α) values. As shown in FIG. 10, with an increase in “add parent threshold” (α) value, the recall score increased and precision score decreased. … The aim was to find an optimum increase in F1 score without drastically affecting the precision score by tuning the (α) and (β) values. The best F1 score was achieved when β=0.5 and α=0.9.”
[0074] “The performance of the classifier models is evaluated based on four metrics—precision, recall, F1 score, and percentage of documents titles with at least one correct prediction tag. … F1 measure is the harmonic mean of precision and recall score. The precision, recall, and F1 scores were calculated for individual tags”
An optimal F1 score is found by tuning parent threshold α. The optimal F1 score is a harmonic mean of precision and recall, therefore recall and precision for a parent label need to be adjusted so that recall and precision can both be high. Since an increase of recall decreases precision, an optimal value of parent threshold α needs to be found that balances both recall and precision, therefore setting a target range (‘threshold’) for precision and recall. The target range for precision being a precision that is high, but low enough not to decrease recall. The target range for recall being a recall that is high, but low enough not to decrease precision.).
The following are the references relied upon in the rejections below:
Goravar (US 20240079102 A1)
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Patil / Pujari / Burns / Gopalakrishnan / Goravar.
Regarding Claim 10, the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan teaches the computer-implemented method of claim 1, however the combination does not teach receiving a query and searching for tagged event data which is taught by Goravar:
further comprising: receiving a query that comprises a given event indicator (Goravar discloses a user may enter (‘query’) desired entities (‘given event indicators’), “the caregiver may enter a set of desired entities into the patient summary system, and the patient summary system may enter each of the desired entities into a respective entity recognition model. Outputs of the respective entity recognition models may be aggregated and refined as described above, to generate the labeled text content. The instances of the desired entities in the labeled text content may be assembled into the data structure” [0085].
Goravar further discloses “a user of the patient summary system may wish to see a listing of all instances of the entity “cancer” in the text content. The patient summary system may request a list of instances of the entity “cancer” found in the text content from the relational database for which a number of instances is greater than 0” [0087].);
searching the tagged event data in the hardware storage device (Goravar discloses instances of desired entities are searched for in labeled text content (‘tagged event data’), “Outputs of the respective entity recognition models may be aggregated and refined as described above, to generate the labeled text content. The instances of the desired entities in the labeled text content may be assembled into the data structure … the patient summary system may search for the instances of the desired entities in the data structure, and may generate the patient summary in accordance with the desired format, based at least partially on data retrieved from the data structure. Because the data structure may be searched more quickly and efficiently than the labeled text content, a speed with which the patient summary may be generated may be increased” [0085]. Searching for instances of entities in labeled text content implies stored data which further implies a hardware storage device.);
and displaying a search result of event data that are tagged with the given event indicator (Goravar discloses “method 400 includes generating a summary of the labeled version of the medical report from the aggregated labeled text data outputted by the model, where the summary summarizes patient information related to the one or more desired entities. To generate the summary, the patient summary system may extract instances of the desired entities, which may be identified by labels as described above, and generate text content based on the entities to display to a caregiver. The text content may include, for example, numbers and types of entities and instances included in the medical report, excerpts of labeled text of the medical report, and/or additional patient data relating to the extracted entities” [0084].).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan with the technique disclosed by Goravar to allow users to retrieve relevant data based on labels. By using labels to organize data, users are able to efficiently lookup and retrieve data based on specific criteria or categories, thereby reducing the time and effort required for data retrieval and analysis.
The following are the references relied upon in the rejections below:
Joyce (US 20210263900 A1)
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Patil / Pujari / Burns / Gopalakrishnan / Joyce.
Regarding Claim 11, the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan teaches the computer-implemented method of claim 1, however the combination does not teach creating and storing a new indicator candidate, which is taught by Joyce:
further comprising: creating, by the machine learning engine, a new indicator candidate ([0149] “For a field, the classification module determines whether the field is already associated with a label in the label index. If a field has not yet been labeled, or if no label index exists, the classification module 605 determines that no label was found for the field. If needed, the classification module 605 generates a new label index to populate with semantic labels”
[0151] “the classification module 605 can be updated over multiple iterations using machine learning approaches”);
and storing the new indicator candidate by the hardware storage device ([0150] “the field can be re-classified by the classification module 605 and re-tested by the testing module 606 to confirm that the label is accurate and potentially update the label attributes of that label in the data dictionary database 614. For example, if the testing module 606 finds the existing label to be a poor fit, a new label can be suggested”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan with the technique disclosed by Joyce to automate data labelling. By automating label creation with machine learning, models can generate new labels or improve existing labels to better describe the data being classified, leading to more interpretable and accurate models.
The following are the references relied upon in the rejections below:
Denney (US 20140015855 A1)
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Patil / Pujari / Burns / Gopalakrishnan / Denney.
Regarding Claim 12, the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan teaches the computer-implemented method of claim 1, however the combination does not teach performing k-means clustering analysis, which is taught by Denney:
further comprising: performing k-means clustering analysis according to the one or more features ([0040] “method includes an initialization phase in block 410, where k distinct labeled descriptors are chosen at random as initial cluster centroids. In some embodiments, the k labeled descriptors are selected to increase the diversity of the initial cluster centroids, for example using k-means++. The initial assigning of a selected labeled descriptor as a cluster centroid may be accomplished by randomly selecting one of the labels assigned to a descriptor from the label assignment binary vector of the descriptor. In order to make the label as specific as possible, if a descriptor has multiple labels and the selected label has one or more child labels assigned to the descriptor (as defined by a label hierarchy), then the cluster label may be changed to a selected one of the child labels. In some embodiments, the child label selection is done at random. This process may be repeated until the selected label has no child labels. This process provides a more specific label as the initial cluster label centroid. Care must be taken to ensure that no two identical centroids (descriptors and selected labels) are repeated in the initial clusters” See Figure 4 depicting a method for clustering using k-means.).
Denney teaches using k-means for clustering descriptors (‘features’) is a known method in the art. Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify the combined tagging method of Patil / Pujari / Burns / Gopalakrishnan with the k-means clustering technique disclosed by Denney to increase the interpretability of classification results. By using k-means to produce clusters, classification results are more interpretable which can help detect outliers and identify underlying patterns in the data.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Sengupta et al. (US 20230061731 A1) teaches assigning TF-IDF scores to tokens extracted from unstructured medical encounter text data to predict an annotation for the data.
Srinivasan et al. (US 20230325468 A1) teaches using a machine learning model to predict a tag for an IT incident event by using an encoded feature vector obtained by encoding text fields of event data.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PEDRO J MORALES whose telephone number is (571)272-6106. The examiner can normally be reached 8:30 AM - 6:00 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MIRANDA M HUANG can be reached at (571)270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PEDRO J MORALES/Examiner, Art Unit 2124
/MIRANDA M HUANG/Supervisory Patent Examiner, Art Unit 2124