Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments, see page 7, filed on 07/01/2026, with respect to claims 15 have been fully considered and are persuasive. The claims interpretation under 35 U.S.C 112(f) of claims 15 has been withdrawn and therefore claim 15 does not invoke 35 U.S.C 112(f).
Applicant’s arguments, see pages 7-8, filed on 07/01/2026, with respect to claims 1-15 have been fully considered and are persuasive. The rejection of claims 1-5 under 35 U.S.C. § 112(b), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter regarded as the invention has been withdrawn.
Applicant’s arguments, see page 10, filed on 07/01/2026, with respect to claims 12, have been fully considered and are persuasive. The rejection of claim 12 under 35 U.S.C. 103 as being unpatentable over NARLIKAR et al. (US 20210406601 A1—hereinafter--"NARLIKAR’) in view of MEZAOUI et al. (US 20210034812 A1 –hereinafter—"MEZAOUI”) has been withdrawn.
Applicant's arguments filed on 07/01/2026 with respect to claims 1-11 and 13-15 have been fully considered but they are not persuasive. The applicant argues that the prior arts of record, Narlikar and Mezaoui, is silent regarding predicting and/or generating labels for abstains or abstain cases of labeling function outputs. However, the examiner respectfully disagrees with the applicant’s argument for at least the following reasons.
The prior art NARLIKAR (US 20210406601) is directed to content classification across media formats which is difficult because each media type often needs its own specialized model and curated training set. Manually labeling rich media such as video or multimodal posts is time-consuming, expensive, and inconsistent [0001, 0038]. It also identifies a “modality gap,” where new content types do not directly map to older labeled examples, making transfer learning and weak supervision harder ([0044, 0064]). NARLIKAR aims to bridge that gap and reduce the need to build separate pipelines from scratch. The application describes a system that classifies content. The system first uses pre-existing classifiers for known media formats to pull out useful features from a new content item. It then combines those features into a shared feature space. Next, it applies simple content rules and uses previously known labeled data to generate training data. That training data is used to train a discrimination model. The trained model can then assign a content policy, such as whether the content should be blocked in a given context. The approach is intended to work even when the new media type has little or no manually labeled training data. The specification emphasizes weak supervision, meaning the system can create labels automatically from rules, statistics, and cross-modal feature matching. It also describes using OCR to read text from images or video frames. This reduces manual labeling effort and speeds development of classifiers for new modalities ([0002-0003], [0039-0041]).
In addition, MEZAOUI (US 20210034812 ) discusses by large volumes of text, especially medical publications, which are difficult to organize and classify. It notes that multi-label classification is harder than single-label classification because one sentence may belong to several categories at once. There is also lack of enough hand-labeled full-text training data for biomedical text. The application seeks a more accurate and scalable way to classify sentences, particularly for Population/Intervention/Outcome-style labeling.
MEZAOUI describes software for tagging a sentence with more than one label at the same time. It is aimed especially at medical text, where a sentence may relate to population, intervention, and outcome all at once. The system first turns the sentence into one or more machine-readable representations, such as embeddings made by BERT or Bio-BERT. It then runs the sentence through a classifier, such as a neural network, to get a probability for each possible label. A separate text-feature score is also computed from the sentence, such as a quantitative-information score or an average TF-IDF score. The system can combine the classifier probabilities with the text-feature score to produce a final output probability for each label. In some versions, a boosting model such as LightGBM performs that combination. In some versions, two representations are used and classified separately, then fused together with the text features. The specification also describes training the classifier using weak supervision and soft labels generated from labelling functions. The goal is to improve sentence classification accuracy, especially for biomedical full-text documents.
The teachings of NARLIKAR in view MEZAOUI combination as disclosed above are in the same filed of endeavor in an effort to teach “predicting and/or generating labels for abstains or abstain cases of labeling function outputs” in (NARLIKAR in [0087-0091] Once common features across data modalities are generated, a set of labeled examples for model training can be curated. One approach to do so can include directly training a model with the labeled data (e.g., labelled categories 128 described herein below, etc.) of existing modalities using the shared features (e.g., extracted features 124, described herein below). However, this technique may be inefficient in certain circumstances. Instead, leveraging existing modalities to generate training data in the target modality without additional manual labeling is proposed as a solution. This can be achieved via weak supervision (WS), which can additionally allow use of features unavailable at deployment time to curate training data. First, an introduction to WS is provided. Techniques to use a common feature space overcome three challenges in using WS for cross-modal adaptation are described. WS can utilize cheap yet noisy labels to curate a training dataset. The techniques described herein present a framework for WS where they generate labeling functions (LFs) are generated that programmatically label groups of data points. To label a set of unlabeled data points, X, unlike a classic, time-consuming labeling pipeline including manual sampling and labeling of individual data points in X, the pipeline described herein proceeds as follows:
1. Develop LFs. Small, labeled development dataset are used to create LFs. LFs can be functions that take a data point and all related features as input, and output a label or abstain (e.g., in a binary setting, an LF returns positive, negative, or abstain). In an example implementation, a sample LF may be: if a post contains excessive profanity it is harmful speech, else abstain. As in the example, while these LFs need not be perfect, the systems described herein can use both high precision and high recall LFs that each perform better than random.
2. Programmatically apply generated LFs to X. Unlike other labeling pipelines, X can be very large as labels are not human-generated. The techniques described herein can be performed on data points in X that have LFs that return labels instead of abstaining, (e.g., have high coverage, etc.).
3. Learn probabilistic labels from Step 2. The systems and methods described herein can use a generative model to estimate each LF's accuracy by evaluating correlations between them when applied to X. The estimated accuracies can be used to return a weighted combination of the weak labels applied to each data point (e.g., probabilistic labels, etc.).
MEZAOUI further discloses predicting and/or generating labels for abstains or abstain cases of labeling function outputs under consideration of data point correlations and/or similarities ([0087] To soft label the plurality of the full-text documents the processor may be to use at least one labelling function to label at least a given portion of each of the full-text documents, for each of the full-text documents the labelling function to: generate one of a set of possible outputs comprising positive, abstain, and negative in relation to associating the given portion with a given label; and generate the one of the set of possible outputs using a frequency-based approach comprising assessing the given portion in relation to at least another portion of the full-text document. [0253] Snorkel is based on the principle of modelling votes from labelling functions as a noisy signal about the true labels. The model is generative and takes into account agreement and correlation between labelling functions, which labelling functions are based on different heuristics. A true class label is modeled as a latent variable and the predicted label is obtained in a probabilistic form (i.e. as a soft label). [0256] UMLS and its metamap tool may be used to automatically extract concepts from medical corpora and, based on heuristic rules, create labelling functions. In some examples, a labelling function may accept as input a candidate object and either output a label or abstain. The set of possible outputs of a specific labelling function may be expanded to include, {positive(+1), abstain(0), negative(−1)} for each of the following classes, population, intervention, and outcome. It is contemplated that similar labeling functions with expanded output sets may also be applied in classification tasks other than medical PIO classification.
Therefore, as presented and discussed above, the applicant’s argument are not persuasive to overcome the prior arts in record and place the claim 1-11 and 13-15 in a better condition for allowance.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-11, 13-17 are rejected under 35 U.S.C. 103 as being unpatentable over NARLIKAR et al. (US 20210406601 A1—hereinafter--"NARLIKAR’) in view of MEZAOUI et al. (US 20210034812 A1 –hereinafter—"MEZAOUI”).
As per claim 1: NARLIKAR discloses a method for
operating a machine learning (ML) system by means of a data processing system,
wherein original data points of a data set are labeled by the data processing system,
(Figure 1: Data Processing Systems 105: [0002] Automatically identify and classify content having new or complex media formats or modalities using weakly supervised machine-learning techniques. Automatically correlate and label previously unlabeled content across complex formats applied to video content, audio content, text content, or any combination. Using feature extraction models, various features can be identified and extracted from content making up the new content modality. A content modality may include portions of video, text, images, or other information. By leveraging existing classifiers for well-known data modalities, this technical solution can identify, extract, and classify features of unknown or new modalities (e.g., combination of media formats, such as a social media post having video, text, images, user information, metadata, etc.). The extracted features can be correlated with previously identified and classified content with similar features to automatically curate and label and training data. This training data can then be used to train a classifier for the new content modality automatically, without using manual labelling processes or manual curation of training data).
the method comprising:
providing the data set and a set of labeling functions for the original data points ([0053] To mitigate the cost of obtaining manually labelled training data, weak supervision systems (e.g., the technical solutions described herein, etc.) are leveraged, and can use labeling functions (LFs) to programmatically label groups of data points);
applying the labeling functions to the original data points for providing a corresponding output of the labeling functions ([0060] This technical solution describes a pipeline that can overcome the challenges of using weak supervision for cross-modal adaptation by automatically generating labeling functions up that are much faster than a domain expert, who must divide the task into days or weeks. Increased performance with respect to coverage and F1 score are obtained by the techniques described herein) the output comprising
labeled data points and labeling function outputs corresponding to each data point ([0071-0072] Organizational resources that process data points of both modalities are identified and applied, and values in a space common between one or more modalities (as in FIG. 2) are output. These common features can be the foundation to using data across modalities in subsequent steps. This step addresses the production challenge of processing rich data modalities, and the cross-modal challenge of bridging the modality gap. Training Data Curation. Labels for the new, unlabeled data modality can be automatically generated to develop a training dataset. To do so, weak supervision can be performed using methods for automatic labeling function creation via frequent item set mining and label propagation that leverage the shared features from the first step. This step can address the production challenge of leveraging diverse information sources, and the cross-modal challenge of decreasing labeling time for rich modalities);
processing at least a part of the output for learning correlations and/or similarities between labeled data points and original data points ([0057] Model Training: combine data and label sources. Given the common features, this technical solution describes leveraging multi-modal techniques for model training that can combine inputs from multiple data and label sources (e.g., data from new and existing modalities, manually generated labels, and labels from weak supervision, etc.). At least three such techniques for combining the features for model training are described: concatenating the features directly, concatenating embeddings independently learned for each data modality, and projecting the new modality to an embedding learned using existing modalities. Combining label and data modalities can improve end-modal performance in comparison to using any modality in isolation, and feature concatenation can outperform the alternatives. [0088] WS can utilize cheap yet noisy labels to curate a training dataset. The techniques described herein present a framework for WS where generate labeling functions (LFs) are generated that programmatically label groups of data points. To label a set of unlabeled data points, X, unlike a classic, time-consuming labeling pipeline including manual sampling and labeling of individual data points in X, the pipeline [0091] 3. Learn probabilistic labels from Step 2. The systems and methods described herein can use a generative model to estimate each LF's accuracy by evaluating correlations between them when applied to X. The estimated accuracies can be used to return a weighted combination of the weak labels applied to each data point (e.g., probabilistic labels, etc.).
NARLIKAR does not explicitly disclose predicting and/or generating labels for abstains or abstain cases of labeling function outputs under consideration of data point correlations and/or similarities. MEZAOUI, in analogous art however, discloses predicting and/or generating labels for abstains or abstain cases of labeling function outputs under consideration of data point correlations and/or similarities ([0087] To soft label the plurality of the full-text documents the processor may be to use at least one labelling function to label at least a given portion of each of the full-text documents, for each of the full-text documents the labelling function to: generate one of a set of possible outputs comprising positive, abstain, and negative in relation to associating the given portion with a given label; and generate the one of the set of possible outputs using a frequency-based approach comprising assessing the given portion in relation to at least another portion of the full-text document. [0253] Snorkel is based on the principle of modelling votes from labelling functions as a noisy signal about the true labels. The model is generative and takes into account agreement and correlation between labelling functions, which labelling functions are based on different heuristics. A true class label is modeled as a latent variable and the predicted label is obtained in a probabilistic form (i.e. as a soft label). [0256] UMLS and its metamap tool may be used to automatically extract concepts from medical corpora and, based on heuristic rules, create labelling functions. In some examples, a labelling function may accept as input a candidate object and either output a label or abstain. The set of possible outputs of a specific labelling function may be expanded to include, {positive(+1), abstain(0), negative(−1)} for each of the following classes, population, intervention, and outcome. It is contemplated that similar labeling functions with expanded output sets may also be applied in classification tasks other than medical PIO classification. Therefore, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the invention to modify the claimed limitations of the labeling function disclosed by NARLIKAR to include predicting and/or generating labels for abstains or abstain cases of labeling function outputs under consideration of data point correlations and/or similarities. This modification would have been obvious because a person having ordinary skill in the art would have been motivated by the desire to provide methods and systems for multi-label classification of a sentence (text) as suggested by MEZAOUI ([0003-0004]).
As per claim 2: NARLIKAR in view MEZAOUI discloses the method according to claim 1, wherein the data set and the set of labeling functions are provided in a knowledge base (NARLIKAR [0081] Model-Based Services. As stated herein, organizations or systems can have access to classification and data processing services that operate over their existing data modalities. Examples can include: topic models that categorize content (e.g., feature classifiers 122 and content rules 126 described herein below, etc.); motif discovery tools to transform time series to categorical patterns (e.g., feature classifiers 122 and content rules 126 described herein below, etc.); knowledge graph querying tools to extract known and related entities from data points (e.g., feature classifiers 122 and content rules 126 described herein below, etc.). The categories, patterns, or the results of various classifications can be stored in the database 115 as the labelled categories 128, as described herein below in conjunction with FIG. 1).
As per claims 3: NARLIKAR in view MEZAOUI discloses the method according to claim 1, wherein applying the labeling functions comprises labeling of data points programmatically (NARLIKAR [0088] WS can utilize cheap yet noisy labels to curate a training dataset. The techniques described herein present a framework for WS where generate labeling functions (LFs) are generated that programmatically label groups of data points. To label a set of unlabeled data points, X, unlike a classic, time-consuming labeling pipeline including manual sampling and labeling of individual data points in X, the pipeline. [0090] 2. Programmatically apply generated LFs to X. Unlike other labeling pipelines, X can be very large as labels are not human-generated. The techniques described herein can be performed on data points in X that have LFs that return labels instead of abstaining, (e.g., have high coverage, etc.).
As per claim 4: NARLIKAR in view MEZAOUI discloses the method according to claim 1, wherein a matrix of labels of labeled data is generated based on the output of the labeling functions (MEZAOUI [0254] Statistical dependencies characterizing the labelling functions and the corresponding accuracy may be modelled. Another factor to consider is propensity, which quantifies and qualifies the density of a given labelling function, i.e., the ratio of the number of times the labelling function is applicable and outputs a label to the original number of unlabeled data. In order to construct a model, the labelling functions may be applied to the unlabeled data points. This will result in label matrix Λ, where Λ.sub.i,j=Λ.sub.i,j (x.sub.i); here, x.sub.i represents the i.sup.th data point and λ.sub.j is the operator representing the j.sup.th labelling function. The probability density function p.sub.w(Λ, Y) may then be constructed using the three factor types that represent the labelling propensity, accuracy, and pairwise correlations of labelling functions).
As per claim 5: NARLIKAR in view MEZAOUI discloses the method according to claim 4, wherein the matrix is amended and/or completed by adding labels resulting from the predicting and/or generating step (MEZAOUI [0255] A concatenation of these factor tensors for the labelling functions j=1, . . . , n and the pairwise correlations C are defined for a given data point x.sub.i as ϕ.sub.i(Λ, Y). A tensor of weight parameters w∈.sup.2n+|C| is also defined to construct the probability density function: [00003]pw(Λ,Y)=Zw-1.Math.exp(.Math.i=1m.Math.wT.Math.φi(Λ,yi))(7) where Z.sub.w is the normalizing constant. In order to learn the parameter w without access to the true labels Y the negative log marginal likelihood given the observed label matrix is minimized: The trained model is then used to obtain the probabilistic training labels, Ŷ=p.sub.{circumflex over (ω)}(Y/Λ), also referred to as soft labels. This model may also be described as a generative model).
As per claim 6: NARLIKAR in view MEZAOUI discloses the method according to claim 4, wherein a component predicting and/or generating labels for abstains or abstain cases of labeling function outputs under consideration of data point correlations and/or similarities predicts abstains or abstain cases in the matrix, wherein the component is a generative machine learning; (ML) (MEZAOUI [0256] UMLS and its metamap tool may be used to automatically extract concepts from medical corpora and, based on heuristic rules, create labelling functions. In some examples, a labelling function may accept as input a candidate object and either output a label or abstain. The set of possible outputs of a specific labelling function may be expanded to include, {positive(+1), abstain(0), negative(−1)} for each of the following classes, population, intervention, and outcome. It is contemplated that similar labeling functions with expanded output sets may also be applied in classification tasks other than medical PIO classification).
As per claim 7: NARLIKAR in view MEZAOUI discloses the method according to claim 6, wherein abstains or abstain cases in the matrix are replaced with certain values or labels resulting from the predicting and/or generating step by the component (MEZAOUI [0257] The labels positive and abstain are represented by +1 and 0 respectively. In order to construct a labelling function, the correlation of the presence of a concept in a sentence with each PIO class may be taken into account. The labelling function for given class j and concept c may be defined as: inequation 10).
As per claim 8: NARLIKAR in view MEZAOUI discloses the method according to claim 1, wherein similarities between original data points comprise distances or values of distances between original data points (NARLIKAR [0056] Finding Borderline Examples. Weak supervision can mandate high precision and recall LFs that cover a majority of data points. While developing LFs to identify positive examples with high precision can be straightforward, constructing rules to identify borderline positive and negative examples, which are crucial for recall and coverage, may be challenging. In response, the techniques described herein can use label propagation to augment the automatically mined LFs. Label propagation can detect data points in the new modality that may be similar to labeled examples in the old modalities, where similarity can be defined using features in the common feature space. Such techniques can enable the identification of large volumes of negative examples, and more candidate positives than with techniques implementing pure item set mining, thereby improving overall computational performance of such systems.
As per claim 9: NARLIKAR in view MEZAOUI discloses the method according to claim 5, wherein the amended and/or completed matrix is fed to a generative machine learning (ML) that chooses a single label for the data points or for any given data point (MEZAOUI [0120] Each of the classifiers 122 can also be stored with a modality identifier that identifies the type of media content the classifier 122 can classify. Each of the classifiers 122 can be treated as a function or algorithm that can take a type of media content as an input. For example, if the classifier 122 is a convolutional deep neural network model to classify certain features of images, the classifier 122 can take an image, or an extracted frame of video, that is formatted to conform to the input of the classifier 122, as an input value (e.g., an input matrix, an input vector, a normalized matrix or vector, or another input data structure, etc.). One or more of the components of the data processing system 105 can process a media type to conform to the input of a classifier 122. The classifier 122 can include one or more layers, model types (e.g., neural network, logistic regression model, linear regression model, convolutional neural network, deep neural network, long short-term memory models, other types of machine learning or artificial intelligent models that can classify features of content, etc.).
As per claim 10: NARLIKAR in view MEZAOUI discloses the method according to claim 1, wherein chosen single labels are used for training a discriminative model (MEZAOUI [0126] The discrimination model trainer 150 can train a discrimination model using the content item and the determinative training data. The discrimination model can be trained to classify or associate a content item with a content policy. The content item can have a data modality or format for which a classifier may not exist. Media formats or modalities can include video, text, audio, images, instructions for constructing interfaces that are subsequently displayed on a device, or any combination thereof. The discrimination model trainer 150 can construct an input vector using the determinative training data and the content item).
As per claim 11: NARLIKAR in view MEZAOUI discloses the method according to claim 4, wherein a heuristic method or a learning algorithm implements a generative machine learning (ML) for reinforcing labels by a Labeling Functions' Reinforcer, wherein the Labeling Functions' Reinforcer amends and/or completes the matrix before the generative model decides on the final array of labels in the matrix (MEZAOUI [0253] Snorkel is based on the principle of modelling votes from labelling functions as a noisy signal about the true labels. The model is generative and takes into account agreement and correlation between labelling functions, which labelling functions are based on different heuristics).
As per claim 12: NARLIKAR in view MEZAOUI discloses the method according to claim 11, wherein in the heuristic method or learning algorithm and/or in the processing step a gravitation process or a clustering process is used, wherein the gravitation process and the clustering process are based on similarities between not labeled data points or abstains or abstain cases and labeled data points (MEZAOUI [0111] To soft label the plurality of the full-text documents the processor may be to use at least one labelling function to label at least a given portion of each of the full-text documents, for each of the full-text documents the labelling function to: generate one of a set of possible outputs comprising positive, abstain, and negative in relation to associating the given portion with a given label; and generate the one of the set of possible outputs using a frequency-based approach comprising assessing the given portion in relation to at least another portion of the full-text document).
As per claim 13: NARLIKAR in view MEZAOUI discloses the method according to claim 1, wherein the data points are vectors, texts or images (NARLIKAR [0120] Each of the classifiers 122 can be stored with a feature type identifier that identifies the type of feature that the respective classifier 122 can classify. Feature type identifiers can include text strings, index values, or other values that indicate that the respective classifier 122 can classify a certain feature. Each of the classifiers 122 can also be stored with a modality identifier that identifies the type of media content the classifier 122 can classify. Each of the classifiers 122 can be treated as a function or algorithm that can take a type of media content as an input. For example, if the classifier 122 is a convolutional deep neural network model to classify certain features of images, the classifier 122 can take an image, or an extracted frame of video, that is formatted to conform to the input of the classifier 122, as an input value (e.g., an input matrix, an input vector, a normalized matrix or vector, or another input data structure, etc.).
As per claim 14: NARLIKAR in view MEZAOUI discloses the method according to claim 1, wherein the method is used in Internet of Things (IoT) or in healthcare (MEZAOUI [0136] With increases in the pace of human creative activity, the volume of the resulting text records continues to increase. For example, the increasing volumes of medical publications make it increasingly difficult for medical practitioners to stay abreast of the latest developments in medical sciences. In addition, the increasing ability to capture and transcribe voice and video recordings into text records further increasers the volumes of text data to organize and classify).
As per claim 15: Claim 15 is directed to a data processing system for carrying out the method for operating a machine learning (ML), wherein original data points of a data set are labeled by the data processing system, the system having substantially similar corresponding limitation of claim 1 and therefore claim 15 is rejected with the same rationale given above to reject corresponding limitations of claim 1.
As per claim 16: NARLIKAR in view MEZAOUI discloses the method according to claim 1, wherein predicting and/or generating labels for abstains or abstain cases of labeling function outputs comprises learning correlations and similarities between labelled data points and unlabeled data points (MEZAOUI [0256] UMLS and its metamap tool may be used to automatically extract concepts from medical corpora and, based on heuristic rules, create labelling functions. In some examples, a labelling function may accept as input a candidate object and either output a label or abstain. The set of possible outputs of a specific labelling function may be expanded to include, {positive(+1), abstain(0), negative(−1)} for each of the following classes, population, intervention, and outcome. It is contemplated that similar labeling functions with expanded output sets may also be applied in classification tasks other than medical PIO classification).
As per claim 17: NARLIKAR in view MEZAOUI discloses the method according to claim 16, predicting and/or generating labels for abstains or abstain cases of labeling function outputs comprises predicting and generating new and/or latent labels for the unlabeled data points based on data point correlations and/or similarities (MEZAOUI [0253-0254] Snorkel is based on the principle of modelling votes from labelling functions as a noisy signal about the true labels. The model is generative and takes into account agreement and correlation between labelling functions, which labelling functions are based on different heuristics. A true class label is modeled as a latent variable and the predicted label is obtained in a probabilistic form (i.e. as a soft label). Statistical dependencies characterizing the labelling functions and the corresponding accuracy may be modelled. Another factor to consider is propensity, which quantifies and qualifies the density of a given labelling function, i.e., the ratio of the number of times the labelling function is applicable and outputs a label to the original number of unlabeled data. In order to construct a model, the labelling functions may be applied to the unlabeled data points. This will result in label matrix Λ, where Λ.sub.i,j=Λ.sub.i,j (x.sub.i); here, x.sub.i represents the i.sup.th data point and λ.sub.j is the operator representing the j.sup.th labelling function. The probability density function p.sub.w(Λ, Y) may then be constructed using the three factor types that represent the labelling propensity, accuracy, and pairwise correlations of labelling functions).
Allowable Subject Matter
Claim 12 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter (the prior art of record either alone or in combination do not overcome claim 12: wherein in the heuristic method or learning algorithm and/or in the processing step a gravitation process or a clustering process is used, wherein the gravitation process and the clustering process are based on similarities between not labeled data points or abstains or abstain cases and labeled data points).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Rogers et al. ( US 20190147297 A1) discuses implementations of the present disclosure are generally directed to labeling training data for training machine learning (ML) models. More particularly, implementations of the present disclosure are directed to a visual platform for relatively rapid assignment of labels to training data based on ontological classes.
Cirillo et al. (US 20230237321 A1) describes a way for multiple organizations or devices in a federated learning system to improve local machine-learning training without sharing their raw data or their labeling rules. Each node first applies its own labeling functions to its own data to create a matrix of noisy labels and abstentions. A separate similarity-computing party then compares data points from different nodes and produces similarity scores. Using those similarity scores, one node can “borrow” labels from another node’s labeling matrix and fill in some of its own missing labels. The updated label matrix is then condensed into one label per data point, with abstains removed. That cleaned dataset is used to train a discriminative model locally. The system is designed to help the local model converge better before any global federated aggregation. The disclosure emphasizes privacy because neither raw data nor labeling-function code needs to be exchanged. It is presented as useful for settings like smart buildings, industrial IoT, healthcare, and energy systems. The same general workflow can be repeated at different nodes in the federation.
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to TECHANE GERGISO whose telephone number is (571)272-3784. The examiner can normally be reached 9:30am to 6:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, LINGLAN EDWARDS can be reached at (571) 270-5440. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/TECHANE GERGISO/ Primary Examiner, Art Unit 2408