DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This office action is in responsive to communication(s): original application filed on 04/16/2024, said application claims a priority filing date of 06/19/2023. Claims 1-2 pending. Claims 1 and 21-22 are independent.
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because (1) reference character “203” has been used to designate both "step" in ¶ [0041] with FIG. 2 and "determination unit" in ¶¶ [0150] and [0197]; and (2) reference character “307” has been used to designate both "computing unit" in ¶¶ [0132]-[0133], [0143], [0147], [0166], [0177], [0179]-[0181], [0186], [0190], [0192], [0202], and [0205] with FIG. 3 and "obtaining unit" in ¶ [0178]. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because (1) reference characters "203" in ¶¶ [0150] and [0197] and "302" in ¶¶ [0132]-[0133], [0135], [0137], [0139]-[0142], [0144]-[0149], [0152]-[0153], [0156]-[0158], [0168]-[0169], [0183], [0185], [0191], and [0198]-[0199] with FIG. 3 have both been used to designate "determination unit"; (2) reference characters "306" in ¶¶ [0132]-[0133], [0140], [0150], [0154], [0162], [0172], and [0188] with FIG. 3 and "307" in ¶ [0178] have both been used to designate "obtaining unit"; and (3) reference characters "408" in ¶¶ [0212]-[0213] with FIG. 4 and "608" in ¶ [0211] have both been used to designate "storage unit". Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they do not include the following reference sign(s) mentioned in the description: 608 in ¶ [0211]. Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The disclosure is objected to because of the following informalities:
in ¶¶ [0006]-[0008] and [0221]-[0223], "… one or more labeling index values corresponding to the one or more labeling asks …" appears "… one or more labeling index values corresponding to the one or more labeling tasks …";
in ¶ [00150], "… the determination unit 203 is also used for determining a target pre-filter procedure …" appears to be "… the determination unit 302 is also used for determining a target pre-filter procedure …";
in ¶ [0164], it appears that something is missing after "and" at the end of paragraph;
in ¶ [0178], "… the obtaining unit 307 is also used for obtaining task characteristics of the target labeling task …" appears to be "… the obtaining unit 306 is also used for obtaining task characteristics of the target labeling task …";
in ¶ [0197], "… if greater, the determination unit 203 also determines that the first automatic labeling procedure satisfies the first quality evaluation index; if not greater, the determination unit 203 also determines that the first automatic labeling procedure fails to satisfy the first quality evaluation index" appears to be "… if greater, the determination unit 302 also determines that the first automatic labeling procedure satisfies the first quality evaluation index; if not greater, the determination unit 302 also determines that the first automatic labeling procedure fails to satisfy the first quality evaluation index";
in ¶ [0205], "… the processing unit 308 is also used for 8updating the first automatic labeling procedure based on …" appears to be "… the processing unit 308 is also used for updating the first automatic labeling procedure based on …"; and
in ¶ [0211], "… programs loaded in the random-access memory (RAM) 403 from a storage unit 608 …" appears to be "… programs loaded in the random-access memory (RAM) 403 from a storage unit 408 …".
Appropriate correction is required.
Claim Objections
Claims 1-22 are objected to because of the following informalities:
in Claim 1, lines 2-6; Claim 21, lines 5-9; and Claim 22, lines 4-8,"… receiving a data labeling request, the data labeling request including one or more labeling tasks, one or more task types corresponding to the one or more labeling tasks, labeling data, a constraint, one or more labeling index values corresponding to the one or more labeling asks, the labeling data including one or more data items, the constraint including one or more constraint conditions …" appears to be "… receiving a data labeling request, wherein the data labeling request including one or more labeling tasks, one or more task types corresponding to the one or more labeling tasks, labeling data, a constraint, and one or more labeling index values corresponding to the one or more labeling tasks, wherein the labeling data including one or more data items, and the constraint including one or more constraint conditions …";
in Claim 2, lines 1-3, "… wherein determining the labeling procedure based on the data labeling request and at least one selected from a group consisting of the task type, the labeling quality, and the labeling metric comprises …" appears to be "… wherein determining the labeling procedure based on the data labeling request and the at least one selected from a group consisting of the task type, the labeling quality, and the labeling metric comprises …";
in Claim 2, lines 6-12, it is recommended to change "… if …" to "… when …" which provide more definite condition in scope;
in Claim 3, lines 7-10, "… determining whether the first automatic labeling procedure satisfies the first quality evaluation index; if the first automatic labeling procedure satisfies the first quality evaluation index, setting the first automatic labeling procedure as the labeling procedure." appears to be "… determining whether the first automatic labeling procedure satisfies the first quality evaluation index; and when the first automatic labeling procedure satisfies the first quality evaluation index, setting the first automatic labeling procedure as the labeling procedure." ;
in Claim 4, lines 3-4, "… computing one or more first numbers corresponding to one or more labeling tasks by at least …" appears to "… computing one or more first numbers corresponding to the one or more labeling tasks by at least …";
in Claim 4, lines 5-9, "… for each labeling task in the one or more labeling tasks, computing a first number of a rest of the labeling tasks dominated by the each labeling task based on a corresponding labeling index value of the each labeling task, wherein the rest of the labeling tasks are one or more remaining labeling tasks after selecting a labeling task from the one or more labeling tasks …" appears to be "… for each labeling task in the one or more labeling tasks, computing a first number of rest labeling tasks dominated by said each labeling task based on a corresponding labeling index value of said each labeling task, wherein the rest labeling tasks are one or more remaining labeling tasks after selecting a labeling task from the one or more labeling tasks …" according to ¶ [0039] of the specification;
in Claim 4, line 10, "… sorting a priority of one or more labeling tasks based on the one or more first numbers …" appears to be "… sorting a priority of the one or more labeling tasks based on the one or more first numbers …";
in Claim 5, lines 1-2, "… wherein dividing, based on the constraint, the labeling data of the target labeling task into automatic labeling data and manual labeling data comprises …" appears to be "… wherein dividing, based on the constraint, the labeling data of the target labeling task into the automatic labeling data and the manual labeling data comprises …";
in Claim 5, lines 5-8, "… assigning one or more data items … satisfying the constraint as the automatic labeling data; assigning one or more data items … not satisfying the constraint as the manual labeling data." Appears to be "… assigning one or more data items … satisfying the constraint as the automatic labeling data; and assigning one or more data items … not satisfying the constraint as the manual labeling data." (see also 112 Rejections to Claim 5);
in Claim 6, lines 1-2, "… wherein assigning one or more data items … not satisfying the constraint as the manual labeling data comprises …" appears to be "… wherein assigning the one or more data items … not satisfying the constraint as the manual labeling data comprises …" (see also 112 Rejection to Claim 5);
in Claim 6, lines 7-9, "… sending one or more data items satisfying the target pre-filter procedure threshold to a manual labeling data item queue; discarding one or more data items not satisfying the target pre-filter procedure threshold." appears to be "… sending one or more data items satisfying the target pre-filter procedure threshold to a manual labeling data item queue; and discarding one or more data items not satisfying the target pre-filter procedure threshold.";
in Claim 7, lines 3-5, "… wherein each pre-filter procedure in the set of pre-filter procedures corresponds to a task type, and each pre-filter procedure includes discarding negatives in the manual labeling data …" appears to be "… wherein each pre-filter procedure in the set of pre-filter procedures corresponds to a task type, and said each pre-filter procedure includes discarding negatives in the manual labeling data …";
in Claim 7, lines 6-9, "… matching the task type of the target labeling task with one or more task types corresponding to the set of pre-filter procedures to identifying a matched pre-filter procedure in the set of pre-filter procedures; assigning the matched pre-filter procedure as the target pre-filter procedure." appears to be "… matching the task type of the target labeling task with one or more task types corresponding to the set of pre-filter procedures to identifying a matched pre-filter procedure in the set of pre-filter procedures; and assigning the matched pre-filter procedure as the target pre-filter procedure.";
in Claim 8, lines 1-3, "… wherein determining, based on the task type of the target labeling task, whether the existing automatic labeling procedure corresponding to the target labeling task is present comprises …" appears to be "… wherein determining, based on the task type of the target labeling task, whether the existing labeling procedure corresponding to the target labeling task is present comprises …" according to Claim 2;
in Claim 9, lines 1-3, "… wherein determining, based on the task type of the target labeling task, whether an existing automatic labeling procedure corresponding to the target labeling task is present further comprises …" appears to be "… wherein determining, based on the task type of the target labeling task, whether the existing labeling procedure corresponding to the target labeling task is present further comprises …" (see also Claim Objections to Claim 8);
in Claim 9, lines 4-5, it is recommended to change "… if …" to "… when …" which provide more definite condition in scope;
in Claim 10, lines 6-10, "… obtaining a result of labeling the quality inspection data by the existing automatic labeling procedure and a result of manual labeling the quality inspection data; determining whether the existing automatic labeling procedure satisfies the first quality evaluation index based on the result of labeling the quality inspection data by the existing automatic labeling procedure and the result of manual labeling the quality inspection data" appears to be "… obtaining a result of labeling the quality inspection data by the existing labeling procedure and a result of manual labeling the quality inspection data; and determining whether the existing labeling procedure satisfies the first quality evaluation index based on the result of labeling the quality inspection data by the existing labeling procedure and the result of manual labeling the quality inspection data" according to Claim 2;
in Claim 11, lines 3-25, 6 instances of "… the existing automatic labeling procedure …" appears to be "… the existing labeling procedure …" (see also Claim Objections to Claim 10);
in Claims 11, lines 16-25, "… in case that the recall rate, the precision rate, the accuracy rate, the false positive rate and the false negative rate are respectively greater than the recall rate threshold, the precision rate threshold, the accuracy rate threshold, the false positive rate threshold and the false negative rate threshold, determining that … satisfies the first quality evaluation index; in case that at least one of the recall rate, the precision rate, the accuracy rate, the false positive rate and the false negative rate is not greater than the recall rate threshold, the precision rate threshold, the accuracy rate threshold, the false positive rate threshold and the false negative rate threshold, determining that … fails to satisfy the first quality evaluation index." appears to be "… in case that the recall rate, the precision rate, the accuracy rate, the false positive rate and the false negative rate are respectively greater than the recall rate threshold, the precision rate threshold, the accuracy rate threshold, the false positive rate threshold and the false negative rate threshold, determining that … satisfies the first quality evaluation index; and in case that at least one of the recall rate, the precision rate, the accuracy rate, the false positive rate and the false negative rate is not greater than the recall rate threshold, the precision rate threshold, the accuracy rate threshold, the false positive rate threshold and the false negative rate threshold, determining that … fails to satisfy the first quality evaluation index";
in Claim 12, lines 1-4, 3 instances of "… wherein determining whether the existing automatic labeling procedure satisfies the first quality evaluation index comprises: if the existing automatic labeling procedure satisfies the first quality evaluation index, labeling the automatic labeling data with the existing automatic labeling procedure" appears to be "… wherein determining whether the existing labeling procedure satisfies the first quality evaluation index comprises: when the existing labeling procedure satisfies the first quality evaluation index, labeling the automatic labeling data with the existing labeling procedure" according to Claim 2;
in Claim 13, lines 3-8, "… selecting … X feature extraction models and Y classifier models respectively based on the task type of the target labeling task, X and Y being positive integers; obtaining … from permutating and combining any one feature extraction model of the X feature extraction models with any one classifier model of the Y classifier models." appears to be "… selecting … X feature extraction models and Y classifier models respectively based on the task type of the target labeling task, X and Y being positive integers; and obtaining … from permutating and combining any one feature extraction model of the X feature extraction models with any one classifier model of the Y classifier models." (see also 112 Rejections to Claim 13);
in Claim 14, lines 1-2, "… wherein computing labeling quality and labeling metric of each automatic labeling procedure in the set of automatic labeling procedures comprises: …" appears to be "… wherein computing the labeling quality and the labeling metric of said each automatic labeling procedure in the set of automatic labeling procedures comprises: …" (see also 112 Rejections to Claim 3);
in Claim14, lines 5-16, "… labeling the quality inspection data respectively with each automatic labeling procedure in the set of automatic labeling procedures; computing labeling quality of each automatic labeling procedure in the set of the automatic labeling procedures based on a result of labeling the quality inspection data by each automatic labeling procedure in the set of the automatic labeling procedures and the result of manual labeling the quality inspection data; obtaining task characteristics of the target labeling task, the task characteristics including quantity of data items of the automatic labeling data and an average cost for manually labeling a single data item of the automatic labeling data; computing labeling metric of each automatic labeling procedure in the set of the automatic labeling procedures respectively based on the task characteristics and the result of labeling the quality inspection data by each of the automatic labeling procedures" appears to be "… labeling the quality inspection data respectively with said each automatic labeling procedure in the set of automatic labeling procedures; computing the labeling quality of said each automatic labeling procedure in the set of the automatic labeling procedures based on a result of labeling the quality inspection data by said each automatic labeling procedure in the set of the automatic labeling procedures and a result of manual labeling the quality inspection data; obtaining task characteristics of the target labeling task, the task characteristics including quantity of data items of the automatic labeling data and an average cost for manually labeling a single data item of the automatic labeling data; and computing the labeling metric of said each automatic labeling procedure in the set of the automatic labeling procedures respectively based on the task characteristics and the result of labeling the quality inspection data by said each automatic labeling procedure." (see also 112 Rejections to Claim 3);
in Claim 15, lines 1-2, "… wherein computing labeling quality of each automatic labeling procedure in the set of the automatic labeling procedures comprises …" appears to be "… wherein computing the labeling quality of said each automatic labeling procedure in the set of the automatic labeling procedures comprises …" (see also 112 Rejections to Claim 3);
in Claim 15, lines 3-9, "… computing the recall rate, the precision rate, the accuracy rate and a ratio of correctly identified negatives of each automatic labeling procedure in the set of the automatic labeling procedures respectively based on the result of labeling the quality inspection data by each of the automatic labeling procedures and the result of manual labeling the quality inspection data; obtaining labeling quality of each automatic labeling procedure in the set of the automatic labeling procedures through linear weighting of the recall rate, the precision rate, the accuracy rate and the ratio of correctly identified negatives." appears to be "… computing a recall rate, a precision rate, a accuracy rate and a ratio of correctly identified negatives of said each automatic labeling procedure in the set of the automatic labeling procedures respectively based on the result of labeling the quality inspection data by said each automatic labeling procedure and the result of manual labeling the quality inspection data; obtaining the labeling quality of said each automatic labeling procedure in the set of the automatic labeling procedures through linear weighting of the recall rate, the precision rate, the accuracy rate and the ratio of correctly identified negatives." (see also 112 Rejections to Claim 3);
in Claim 16, lines 3-7, "… computing a Pareto dominance relation between any two automatic labeling procedures in the set of the automatic labeling procedures based on the labeling quality and the labeling metric; determining the first automatic labeling procedure based on the Pareto dominance relation." appears to be "… computing a Pareto dominance relation between any two automatic labeling procedures in the set of the automatic labeling procedures based on the labeling quality and the labeling metric; and determining the first automatic labeling procedure based on the Pareto dominance relation.";
in Claim 17, lines 1-12, "… wherein computing Pareto dominance relation between any two automatic labeling procedures in the set of the automatic labeling procedures based on the labeling quality and the labeling metric comprises … determining whether the first labeling metric is smaller than or equal to a labeling metric threshold; if the first labeling metric is smaller than the labeling metric threshold, traversing the set of labeling qualities and the set labeling metrics to calculate Pareto dominance relation between any two automatic labeling procedures in the set of the automatic labeling procedures." appears to be "… wherein computing the Pareto dominance relation between said any two automatic labeling procedures in the set of the automatic labeling procedures based on the labeling quality and the labeling metric comprises … determining whether the first labeling metric is smaller than or equal to a labeling metric threshold; and when the first labeling metric is smaller than the labeling metric threshold, traversing the set of labeling qualities and the set labeling metrics to calculate the Pareto dominance relation between said any two automatic labeling procedures in the set of the automatic labeling procedures.";
in Claim 18, lines 3-10, "… computing one or more second numbers corresponding to the set of automatic labeling tasks by at least: for each automatic labeling procedure in the set of automatic labeling procedures, computing a second number of rest automatic labeling procedures dominated by each automatic labeling procedure in the set of the automatic labeling procedures based on the Pareto dominance relation, where the rest automatic labeling procedures include remaining automatic labeling procedures after selecting any one of the automatic labeling procedure from the set of automatic labeling procedures …" appears to be "… computing one or more second numbers corresponding to the set of automatic labeling procedures by at least: for said each automatic labeling procedure in the set of automatic labeling procedures, computing a second number of rest automatic labeling procedures dominated by said each automatic labeling procedure in the set of the automatic labeling procedures based on the Pareto dominance relation, where the rest automatic labeling procedures include remaining automatic labeling procedures after selecting any one automatic labeling procedure from the set of automatic labeling procedures…";
in Claim 18, lines 11-14, "… sorting a priority of the set of automatic labeling procedures based on the one or more second numbers; determining the first automatic labeling procedure as an automatic labeling procedure with the highest second number in the one or more second numbers." appears to be "… sorting a priority of the set of automatic labeling procedures based on the one or more second numbers; and determining the first automatic labeling procedure as an automatic labeling procedure with the highest second number in the one or more second numbers."
in Claim 19, lines 6-9, "… determining whether the first automatic labeling procedure is greater than a quality control threshold of the first quality evaluation index based on a result of labeling the quality inspection data by the first automatic labeling procedure and the result of manual labeling the quality inspection data …" appears to be "… determining whether the first automatic labeling procedure is greater than a quality control threshold of the first quality evaluation index based on a result of labeling the quality inspection data by the first automatic labeling procedure and a result of manual labeling the quality inspection data …";
in Claim 19, lines 10-15, "… if the first automatic labeling procedure is greater than the quality control threshold, determining that the first automatic labeling procedure satisfies the first quality evaluation index; if the first automatic labeling procedure is not greater than the quality control threshold, determining that the first automatic labeling procedure fails to satisfy the first quality evaluation index." appears to be "… when the first automatic labeling procedure is greater than the quality control threshold, determining that the first automatic labeling procedure satisfies the first quality evaluation index; and when the first automatic labeling procedure is not greater than the quality control threshold, determining that the first automatic labeling procedure fails to satisfy the first quality evaluation index.";
in Claim 20, lines 3-6, "… assigning the automatic labeling data of the target labeling task as unlabeled data; updating the first automatic labeling procedure until a Pareto target value of the first automatic labeling procedure converges, wherein …" appears to be "… assigning the automatic labeling data of the target labeling task as unlabeled data; and updating the first automatic labeling procedure until a Pareto target value of the first automatic labeling procedure converges, wherein …";
in Claim 20, lines 14-18, "… selecting data items from the randomly extracted data … to form extraction data with high contribution value; manually labeling the extraction data with a high contribution value to obtain a corresponding manual labeling result …" appears to be "… selecting data items from the randomly extracted data … to form extraction data with a high contribution value; manually labeling the extraction data with the high contribution value to obtain a corresponding manual labeling result …";
in Claim 20, lines 26-27, "… continuing to perform the extraction and labeling operations based on the unlabeled data updated until the second Pareto target value converges" appears to be "… continuing to perform the extraction and labeling operations based on the unlabeled data updated until the second Pareto optimal target value converges".
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-22 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 1 and 21-22 recite the limitation "… the data labeling request including " in lines 2-10, 5-13, and 4-12 respectively, which rendering these claims indefinite because it is unclear ".
Claims 1 and 21-22 recite the limitation "… a group consisting of the task type, a labeling quality, and a labeling metric …" in line 12, lines 14-15, and 13-14 respectively. There is insufficient antecedent basis for the limitation "the task type" in the claim. For examination purpose, … a group consisting of a task type, a labeling quality, and a labeling metric …" is considered.
Claims 2-20 are rejected for fully incorporating the deficiency of their respective base claims.
Claim 2 recites the limitation "… determining, based on a task type of the target labeling task, whether an existing labeling procedure corresponding to the target labeling task is present …" in lines 4-5, which rendering the claim indefinite because ".
Claims 3 and 8-20 are rejected for fully incorporating the deficiency of their respective base claims.
Claim 3 recites the limitation "… wherein determining, based on the set of automatic labeling procedures, the labeling procedure comprises: computing a labeling quality and a labeling metric of each automatic labeling procedure in the set of automatic labeling procedures; determining a first automatic labeling procedure based on the labeling quality and the labeling metric …" in lines 1-6, which rendering the claim indefinite because ".
Claims 14-20 are rejected for fully incorporating the deficiency of their respective base claims.
Claim 4 recites the limitation "the highest first number" in line 12. There is insufficient antecedent basis for this limitation in the claim.
Claim 5 recites the limitation "… wherein dividing, based on the constraint, the labeling data of the target labeling task into … comprises: deciding whether each data item of the one or more data items in the labeling data satisfies each constraint condition of the one or more constraint conditions of the constraint; assigning one or more data items of the labeling data satisfying the constraint as the automatic labeling data; assigning one or more data items of the labeling data not satisfying the constraint as the manual labeling data" in lines 1-8, which rendering the claim indefinite because ".
Claims 6-7 are rejected for fully incorporating the deficiency of their respective base claims.
Claim 6 recites the limitation "… determining a target pre-filter procedure based on a task type of the target labeling task …" in line 3, which rendering the claim indefinite because ".
Claim 7 is rejected for fully incorporating the deficiency of their respective base claims.
Claim 13 recites the limitation "… wherein obtaining the set of automatic labeling procedures based on the task type of the target labeling task comprises: … obtaining a set of automatic labeling procedures from permutating and combining any one feature extraction model of the X feature extraction models with any one classifier model of the Y classifier models" in lines 1-8, which rendering the claim indefinite because ".
Claim 14 recites the limitation "... extracting quality inspection data from the labeling data, the quality inspection data is a part of the labeling data ..." in lines 3-4, which rendering the claim indefinite because ".
Claim 15 is rejected for fully incorporating the deficiency of their respective base claims.
Claim 18 recites the limitation "the highest first number" in line 1. There is insufficient antecedent basis for this limitation in the claim.
Claim 19 recites the limitation "... extracting quality inspection data from the labeling data, the quality inspection data is a part of the labeling data ..." in lines 3-4, which rendering the claim indefinite because ".
Claim 19 is rejected for fully incorporating the deficiency of their respective base claims.
Claim 20 recites the limitation "… computing a contribution value of respective data items in the randomly extracted data to a multi-objective optimization model … selecting data items from the randomly extracted data based on contribution values of the respective data items in the randomly extracted data …" in lines 10-15, which rendering the claim indefinite because (1) it.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefore, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract idea without significantly more.
Independent Claims 1 and 21-22
Step 1: Claim 1 is a process claim, Claim 21 is a device claim, and Claim 22 is a claim for a non-transitory computer readable storage medium. These claims fall within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) recite(s) additional elements/limitations of .
Step 2B: The claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the additional limitation/element of .
Claim 2
Step 1: Claim 2 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim 3
Step 1: Claim 3 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim 4
Step 1: Claim 4 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim 5
Step 1: Claim 5 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim 6
Step 1: Claim 6 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) .
Step 2B: The claim(s) does/do not include further additional elements that are sufficient to amount to significantly more than the judicial exception because the additional limitation/element of .
Claim 7
Step 1: Claim 7 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) .
Step 2B: The claim(s) does/do not include further additional elements that are sufficient to amount to significantly more than the judicial exception because the additional limitation/element of .
Claim 8
Step 1: Claim 8 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) .
Step 2B: The claim(s) does/do not include further additional elements that are sufficient to amount to significantly more than the judicial exception because the additional limitation/element of .
Claim 9
Step 1: Claim 9 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim 10
Step 1: Claim 10 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) .
Step 2B: The claim(s) does/do not include further additional elements that are sufficient to amount to significantly more than the judicial exception because the additional limitation/element of .
Claim 11
Step 1: Claim 11 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) "determining whether the existing labeling procedure satisfies a first quality evaluation index, wherein the first quality evaluation index comprises a recall rate, a precision rate, an accuracy rate, a false positive rate and a false negative rate", ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim 12
Step 1: Claim 12 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim 13
Step 1: Claim 13 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim 14
Step 1: Claim 14 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) .
Step 2B: The claim(s) does/do not include further additional elements that are sufficient to amount to significantly more than the judicial exception because the additional limitation/element of .
Claim 15
Step 1: Claim 15 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim 16
Step 1: Claim 16 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim 17
Step 1: Claim 17 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) .
Step 2B: The claim(s) does/do not include further additional elements that are sufficient to amount to significantly more than the judicial exception because the additional limitation/element of .
Claim 18
Step 1: Claim 18 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim 19
Step 1: Claim 19 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) ".
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim 20
Step 1: Claim 20 is a process claim which falls within at least one of the four categories of patent eligible subject matter.
Step 2A Prong 1: The claim(s) further recite(s) "
Step 2A Prong 2: This judicial exception is not integrated into a practical application because the claim(s) does/do not further recite(s) additional elements/limitations.
Step 2B: The claim(s) does/do not further include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, none of the additional limitations, taken either alone or combined, amount to significantly more than the abstract idea.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-9, 12, 14-15, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Prendki (US 2022/0277362 A1, pub. date: 09/01/2022), hereinafter in view of Martin et al. (US 2021/0042577 A1, pub, date: 02/11/2021).
Independent Claims 1 and 21-22
Prendki discloses a method for data labeling (Prendki, ¶[0003]: labeling of datasets for use in supervised machine learning (ML) processes), the method comprising:
receiving a data labeling request, the data labeling request including one or more labeling tasks, one or more task types corresponding to the one or more labeling tasks, labeling data, a constraint, one or more labeling index values corresponding to the one or more labeling asks, the labeling data including one or more data items, the constraint including one or more constraint conditions; determining a target labeling task based on the one or more labeling index values and the one or more labeling tasks (Prendki, ¶¶ [0055]-[0057], [0060], [0063], and [0067]-[0110] with FIG. 3: label service providers 314, 316, 318 comprise any of a server computer, a virtual computing instance hosted in a public or private data center or cloud computing service, a mobile computing device, personal computer, or other computer depending on the performance metrics that are desired in the system; such as throughput, response time, or storage capability; user computer 302 is programmed to transmit, to recommendation computer system 306, a user dataset 304 comprising two or more of a dataset to be labeled, instructions for labeling, and a machine learning model; integrate a data preparation system having data curation technology by employing an Active Learning approach; machine learning models that customers provide can include linear classifiers, neural networks, or other ML models; common use cases that customers may work on include: Image classification; Object detection (2D and 3D); Semantic segmentation (2D and 3D); Instance segmentation (2D and 3D); Pose detection; Facial recognition; Event detection (audio and video); Text classification (including but not limited to sentiment analysis and topic modeling); Content moderation (text, image, audio); Audio transcription; Machine translation; Anomaly detection; Regression. Types of models include: Deep learning models (all architectures), and these models can be pre-trained/off the shelf or custom-made by the user), including but not limited to CNN; RNN; LSTM; MLP; SOM; DBM; GAN; Classical ML models, such as (not limited to) Random forest; Logistic regression; XGBoost; determine the minimum criteria required for handling the labeling tasks requested by the customers; determining the type of the customers' data (e.g., image, video, text, etc.) and the type of labeling task; for task filtering, customer, when creating a labeling task request, may be asked to provide the data type and task type; every labeling task created in association to the customer's project may then automatically be recognized with the matching data type and task type; the various types of tasks may include (a) for Computer Vision (image-based data), Image classification, Object detection, Object segmentation, Instance segmentation, Image captioning, Object counting ( extrapolation), Object tracking, and Event detection; (b) for Natural Language Processing (text-based data), Text classification, Entity extraction, Search relevance (by ranking or scoring), Similarity matching, Text summarization, Question-Answering, and Translation; (c) for Audio Data, Classification, Transcription (speech-to-text), and Event detection; and (d) for Numerical Data, Medical diagnosis (e.g., diagnosis based on blood pressure, temperature, etc.) and Financial analysis (e.g., stock performances, etc.); ¶¶ [0111]-[0112] with FIG. 5A: FIG. 5A illustrates a computer display device that is displaying an example of a graphical user interface with which a customer can provide information about its data labeling requirements; in the example of FIG. SA, a "Data & Task Types" link 505 has been selected and has caused displaying panel 507 having displays and widgets associated with specifying customer data and task types; Panel 507 comprises a data type prompt 506 to which a client device can respond with input to select one of a plurality of data type icons 508, for example, specifying image data, text data, or numeric data; Panel 507 comprises a data task prompt 510 to which the client device can respond with input to select one of a plurality of task checkboxes 512 to specify a particular kind of data processing task that the customer needs; ¶ [0170]: customers may choose their prioritization criteria, including placing certain time constraints on the labeling partners, using the customer platform; ¶¶ [0178]-[0180] with FIG. 6A: with the self-serve process, the user creates a labeling task on the system, by adding info such as: the amount of data that needs to be labeled (that info can be generated by the data curation engine in the case of a dynamic labeling task); the type of data (text, image, video, audio, ... ); the task type (object detection, classification, regression, segmentation, . . . ); the budget ($); their timeline; and their quality (labeling accuracy) requirements; with the bidding process, the user creates a labeling task on the system, by adding info such as: the amount of data that needs to be labeled (that info can be generated by the data curation engine in the case of a dynamic labeling task); the type of data (text, image, video, audio, ... ); the task type ( object detection, classification, regression, segmentation, ... ); with the reverse-bidding process, the user creates a labeling task on the system, by adding info such as: the amount of data that needs to be labeled (that info can be generated by the data curation engine in the case of a dynamic labeling task); the type of data (text, image, video, audio, ... ); the task type (object detection, classification, regression, segmentation, . . . ); relative importance of price, time and accuracy (done through a point allocation system using a front-end interface; the user will allocate 12 points across 3 priorities); ¶ [0191]: tell us about your project, expectations, and time and budgetary constraints; ¶ [0206]: labeling instructions are customer-provided instructions that define the scope of the labeling task' labeling instructions are intended for annotators and provide guidance to how the data should be labeled, e.g., by providing examples of correctly and incorrectly labeled data, covering edges cases and how they should be handled, describing best practices, and providing any other information that may be relevant for the labeling task; labeling instructions are typically tailored based on the input from both the customer and the labeling partner by weighing the customer's expectations with the capabilities of the labeling partner; the universal labeling instructions are designed to be universally compatible with all of the labeling partners within the computer-implemented labeling marketplace by considering each of the labeling partners' capabilities at both the provider-level and annotator-level and the tooling available to the partners; the universal labeling instructions may include a required section and an optional section; ¶ [0218] with 101 in FIG. 1 and 310 in FIG. 3: at step 101, a customer desiring a labeling service uploads the data and labeling instructions to the system; based on the uploaded information, the system determines the type of data provided by the customer (e.g., image, video, text, etc.) and the type of the requested task (e.g., object detection, classification, object segmentation, etc.); once the types of data and task are determined, the system determines the minimum criteria necessary to handle the task requested by the customer, which may include having access to particular tooling (e.g., hardware or software), being capable of understanding certain languages, having certain security clearance, or having a sufficient number of available annotators for the task; to provide labeling services to a customer, the labeling partner needs to fulfill the following expectations: (a) possess a task force capable/trained to manage the particular type of task; (b) possess the proper annotation tool for the said task; be compliant to the level required by the user; (c) be compliant to the level required by the user; and (d) have at least one annotator available immediately; provider metrics database 310 (FIG. 3) containing values concerning the task force, annotation tool, and compliance capabilities of each labeling partner, which the partner provides at the time of enrollment or subscription, and maintains over time through the labeling partner portal; use provider metrics database 310 to retrieve a list of valid candidate partners after taking the input of the user from the frontend; ¶¶ [0226]-[0228] with FIGS. 1-2: at step 201, a customer uploads to the system labeling instructions, data to be labeled, and a model to be trained using the data; at step 202, the customer decides a labeling approach (e.g., ML-based approach, manual approach, or self-labeling approach); at step 203, the curation process begins, and at step 204, the customer's model and data are sent to the data curation engine; the system's data preparation and data curation services employ an Active Learning approach where portions of the data are prioritized so the portions that are most effective in training the model are labeled first; the curation engine analyzes the data and identifies a first batch of data that it believes will be effective in training the data; the first batch of data is provided to a labeling partner in a similar fashion as the processes described in FIG. 1);
choosing (Prendki, ¶¶ [0191]-[0200] with FIG. 8: Hybrid Labeling Marketplace: find awesome labeling partners that are available immediately; process flow: 1. tell us about your project, expectations, and time and budgetary constraints; 2. choose between the recommended auto-labeling and manual labeling options; Mix-and-Match; 3. you get a precise estimate of the cost, time and accuracy; once the task launches, your data is immediately being labeled; 4. use the human-in-the-loop module to find mislabeled data and send it for revision either to the same partner or a different partner; human-in-the-loop module: visualize, audit, and fix your labels and measure labeling quality reliably; 1. upload your existing labels or select a labeling task after the labels are sent back by your labeling partner; 2. view label anomalies or query data using your own validation criteria; no need to ever review an entire dataset anymore; 3. fix faulty labels yourself with an annotation tool or send them back to either the same or a different labeling partner; 4. as you perfect your labels, they are automatically versioned, so that you can revert to an older version if needed; ¶ [0209]: if the system determines that a labeling task includes an amount of data that is too large for any one of the labeling partners in the selection pool, the system may assign various portions of the task to multiple labeling partners; the system may also assign a labeling task to several labeling partners if it determines that diversity among annotators may produce higher labeling quality);
determining a labeling procedure based on the data labeling request and at least one selected from a group consisting of the task type, a labeling quality, and a labeling metric (Prendki, ¶ [0050]: efficiently compare the capabilities, availability, performance, and costs of various labeling providers; a computer-implemented labeling marketplace, through which a suitable labeling partner (i.e., labeling provider partnered with the computer-implemented labeling marketplace) will be automatically selected, based both on historical and real-time metrics, to handle the task requested by a customer; provide a short-list of labeling partners that are similarly determined to be suitable for handling the task; such selection processes may involve eliminating from the selection pool labeling partners that are not capable of handling the requested task; determining real-time availability of the labeling partners, and evaluating the labeling partners' performance history, costs, and various other factors that may be relevant to delivering on the expectations of the customers; incorporate a quality control process, in which the tasks completed by the labeling partners may be evaluated; ¶¶ [0056]-[0057], [0059], [0065], [0067],and [0110] with FIG. 3: the provider metrics database 310 is programmed with a data schema capable of storing a plurality of performance values or metrics for each of the label service providers 314, 316, 318; recommendation computer system 306 is programmed using the real-time evaluation instructions 308 to query the provider metrics database 310 and to select, automatically and based upon programmed recommendation or evaluation algorithms, one or more of the label service providers 314, 316, 318 that is optimal to perform one or more labeling tasks on the user dataset; provides labeling services to customers by finding a labeling partner that is suitable to handle the tasks requested by the customers; a customer may rely on the system to automatically select the labeling partner that is most likely to provide the highest labeling quality in a timely manner, or alternatively, the system may provide the customer with a short-list of labeling partners determined to be suitable for handling the requested task; many labeling tasks require specialized knowledge in particular technologies or languages (e.g., computer vision, translation) or access to a specific kind of tools (e.g., software or hardware for Light Detection and Ranging, or LiDAR); some of these specialized or smaller labeling providers are able to provide superior labeling accuracy and faster tum-around time for certain types of tasks than their larger counterparts because they employ annotators that are more capable, knowledgeable, and skilled in handling those types of tasks; allowing customers to rely on a computer-implemented labeling marketplace to find the labeling provider that is most likely to produce the highest labeling quality and in a timely-manner, regardless of the size of the labeling providers and for whatever types of labeling tasks; labeling partners unable to perform the labeling task because of an inability to perform tasks with the data type and task type requested may be eliminated from consideration for the labeling task; the minimum criteria may be determined during one of the initial phases of the system's processes so any labeling partners that are not capable of handling the requested tasks could be eliminated from the selection pool; ¶¶ [0114]-[0154]: collect various information of the labeling partners, which may be used by the optimization system to find the most suitable labeling partner for a customer-requested task; examples of such information that the system collects of the labeling partners may include the following: list of the types of tasks that the partner can deliver on (depends on tooling and labor workforce); available clearance; annotator information (e.g., type of tasks annotators are capable of working on); historical performance of the annotators (e.g., speed and accuracy); desired workload of annotators; and real-time availability of annotators; compute and maintain various metrics of the labeling providers; examples of such metrics may include a) average labeling quality (e.g., accuracy) for various types of tasks; b) average labeling speed for various types of tasks; c) typical availability; d) average price, and deviation from standard price; e) time since last offered task; f) time of idle (e.g., no assigned task); g) time taken to resolve labeling issues identified by the quality control process; and number of iterations required to solve a labeling issue; h) response time to the system's inquiry/notifications; and i) acceptance/rejection/ignore statistics of labeling tasks; utilize a multi-objective optimization process to find a labeling partner that is most likely to produce the highest quality of work; examples of the factors considered during the optimization process include, but are not limited to, the following: a) turn-around time (e.g., on average how long does the labeling partner need to complete and return a task); b) labeling quality (e.g., how accurate is the labeled data); c) average cost (e.g., the average cost charged by the labeling partner for various types of tasks); d) revenue (e.g., how much did the labeling partner earn for a given period of time); e) number of, and type of, tasks previously offered to the labeling partner; f) number of labeling partners within the computer-implemented labeling marketplace; and g) customer's preferences/criteria; the optimization process may be programmed using any of: the Dantzig simplex algorithm, extensions or variants thereof; combinatorial algorithms; or quantum optimization algorithms; ¶ [0168]: identify a labeling partner that is most suitable for handling the customer's task by considering, among others, real-time availability of the annotators employed by the various labeling partners; allows the system to track each of their annotator's computer activity, or to provide some other indicators to verify whether the annotators are available to receive a labeling task; ¶¶ [0169]-[0173] with FIG. 5F: to provide customers with consistent and efficient turn-around time for the delivery of the requested tasks, the system may place certain time constraints on the labeling partners; the system may provide the labeling partner a limited amount of time to complete any task accepted by the labeling partner, which time period may be determined based on the historical data of the labeling partner or other labeling partners (e.g., average amount of time needed to complete a task of a similar type); the customer may be shown a ranked list of labeling partners and be requested to select its top three choices; the first selected labeling partner may be offered the labeling task; if the first selected labeling partner fails to respond in time, the system may automatically route the request to the labeling partner that was the customer's second top choice; if all three of the customer's initial selections decline the request or fail to timely respond, the system may notify the customer of this failed status and may prompt the customer to continue to select additional labeling partners or to make adjustments to its labeling methodology until a suitable labeling partner accepts the offer for the labeling task; FIG. 5F illustrates a computer display device that is displaying an example of a graphical user interface with which a customer can select one or more labeling partners and view recommendations of labeling partners; ¶¶ [0176]-[0183] with FIG. 6B: select a labeling partner that is suitable for the customer's requested task in a sequential manner, such that the task is offered to only one labeling partner at a time; the system may re-evaluate all of the labeling partners within the selection pool, e.g., in situations where enough time has passed, and real-time metrics associated with the labeling partners have substantially changed; select a labeling partner that is suitable for the customer's requested task by employing a self-serve process, bidding process, or reverse-bidding process; if, e.g., a customer prioritizes cost as an important criteria for the task, the task may be offered to several labeling partners at the same time; in such a case, the labeling partner that provides the lowest cost for the task may be selected; alternatively, the system may provide to the customer a list of labeling providers determined to be suitable for the task, along with the costs associated with each of the providers; the system (which has historical data for price, time and accuracy from previous jobs, on a use case-by-use case basis), is programmed to compute the z-score of each valid labeling partner for price, time and accuracy; each of the z-scores is weighted with the number of points allocated to the feature (for price and time, a -1 factor is added, since lower price or time is better); the system is programmed to compute the sum and to return a ranked list of recommendations based on the sum values; FIG. 6B illustrates an example of a graphical user interface that can be programmed as presentation output of an ordered set of recommended labeling partners; provide task-level feedback that provides information about why a specific task was not offered to a labeling partner (e.g., insufficient availability of annotators, lack of tooling, etc.) or aggregated feedback that provides information about why specific types of tasks, or tasks in general, are not offered to the labeling partner (e.g., statistical insights about subpar turn-around time or labeling accuracy); ¶¶ [0208]-[0209]: select a labeling partner to handle the labeling at an experiment-level (i.e., the task in its entirety), which may involve one or more loops of labeling; this will provide consistency in the labeled data and minimize the inefficiencies that come from transitioning the task between several labeling partners; select several labeling partners to handle the task, each labeling partner being assigned to a particular loop of the task; e.g., if the system determines that a task can be partitioned into several loops and some of the loops are better suitable for one labeling partner while other loops are better suitable for another labeling partner, the system may assign the task to both labeling partners; ¶¶ [0218]-[0221] with FIG. 1: to provide labeling services to a customer, the labeling partner needs to fulfill the following expectations: (a) possess a task force capable/trained to manage the particular type of task; (b) possess the proper annotation tool for the said task; be compliant to the level required by the user; (c) be compliant to the level required by the user; and (d) have at least one annotator available immediately; at step 102, the system applies the minimum criteria to each of the pre-selected labeling partners within the selection pool and eliminates any labeling partners that do not meet the minimum criteria, i.e., that are not capable of handling the task; at steps 103 and 104, the system applies an optimization process that evaluates real-time and historical metrics of all the labeling partners remaining in the selection pool; at step 103, check the real-time availability of the labeling partners, which may include checking real-time availability of the annotators employed by each of the labeling partners; at step 104, evaluates the historical metrics along with the real-time metrics; examples of the historical metrics evaluated by the system may include each of the labeling partners' average labeling accuracy, average turn-around time, and average pricing, both at an overall level and task-specific level (e.g., by grouping together similar types of tasks); based on the real-time and historical metrics, identify a labeling partner that is the most likely to produce the highest quality of work and in a timely-manner; at step 109, if a labeling partner either refuses the task or does not reply within a set amount of time, remove the labeling partner from the selection pool and identifies another labeling partner that is suitable for handling the task; identify the subsequent labeling partner by re-evaluating the real-time and historical metrics of the labeling partners left in the selection pool; alternatively, identify the subsequent labeling partner based on the previous evaluation, e.g., in situations where insignificant amount of time has passed since the previous evaluation; ¶ [0229] with FIGS. 1-2: the steps of identifying/selecting an optimal labeling partner (i.e., steps 102-108 within brackets 260) may be skipped for loops n> 1, if the system determines that the same labeling partner should be used throughout the experiment; alternatively, the steps within brackets 260 may be kept in place if the system determines that different labeling partners should be used within the same experiment); and
labeling the automatic labeling data using the labeling procedure (Prendki, ¶¶ [0057], [0060], [0063]-[0064], and [0174] with FIG. 3: recommendation computer system 306 is programmed to transmit the user dataset to the selected label service provider; to receive a labeled user dataset 312 from the selected label service; provider; to evaluate the performance of the label service provider and update the database; and to transmit the labeled user dataset to the user computer; integrate a data preparation system having data curation technology by employing an Active Leaming approach; a data labeling process integrating a data preparation and data curation processes is called an "experiment" and is often executed in iterations, or "loops"; the Active Learning approach disclosed herein is a form of a semi-supervised, iterative machine learning approach of prioritizing portions of the data so the portions that are most effective in training the machine learning model are labeled first; provide the customer's data to a labeling partner in small batches; when the labeling partner finishes labeling the first batch of data, the system may train a customer-provided model with the labeled data, then evaluate the model's performance; if the system determines that the model has been sufficiently trained, the trained model and the labeled data may be provided to the customer; if the system determines that the model has not been sufficiently trained, the process may be repeated by providing the labeling provider with the next batch of data along with all the previous batches of data; if the system determines that the model's performance is acceptable, the experiment ends, and the customer is provided with the labeled data and trained model; alternatively, the system may end the experiment if it determines that the model has reached its optimal performance; the system may also end the experiment if the customer's budget has run out; ¶¶ [0186]-[0189]: verify the quality of the labeling tasks completed by the labeling partners by evaluating the completed tasks; in ML-driven data curation, which is active learning-based; data is selected dynamically during a training process; the training process ends when the user reaches their budget or the remaining data contains no new relevant information; ¶¶ [0221]-[0223] with 110-112 in FIG. 1: if the task offered to the labeling provider is accepted, at step 110, provide the customer's data and labeling instructions to the labeling partner; at step 111, after completing the labeling task, the labeling partner uploads the labeled data into a corresponding storage database, allowing the system to gain access to it; then, at step 112, evaluate the quality of the labeling partner's work using the system's internal tools; if the quality of the labeling partner's work is determined as not acceptable (e.g., accuracy issues), provide the labeling partner an opportunity to remedy the issues; however, if determine that the labeling partner is unable to address the issues or the labeling partner notifies to the system that it is unable to address the issues, provide the task to another labeling partner; at step 113, if determine that the quality of the labeling partner's work is acceptable, the labeling task is deemed complete and the labeled data is provided back to the customer (e.g., to a data storage dedicated to the customer); store the labeled data in a historical database and index the data with a unique identification assigned to the labeling partner; the system may also index the data based on the type of data labeled, the type of the labeling task, and the quality of the labeled data (e.g., labeling accuracy, turn-around time); the historical metrics evaluated in step 104 may correspond to the information stored in the historical database; ¶¶ [0228]-[0229] with FIG. 2: once the labeling partner completes labeling the first batch of data, the labeled data is used to train the customer's model, as described in step 206; then, the model's performance is evaluated; at steps 207 and 208, if the system determines that the model's performance is acceptable, the system deems the model as sufficiently trained and provides the trained model and labeled data back to the customer (i.e., the experiment ends); the experiment may also end if the system determines that the customer's budget has ran out or if the model has reached its optimal performance; at step 205, if the system determines that the model's performance is unacceptable, the next loop of the experiment begins; going back to step 204, the data curation engine identifies the next batch of data and provides this batch to the labeling partner along with all of the previous batches of data; once the labeling partner completes labeling a batch of the data, the experiment continues until the model's performance reaches an acceptable level, the budget is reached, or optimal performance has been achieved because the remaining data is determined as redundant or irrelevant for training the model).
Prendki further discloses an electronic device (Prendki, ¶¶ [0231]-[0233] with 400 in FIG. 4: the techniques described herein are implemented by at least one computing device; a computer system 400 for implementing the disclosed technologies), comprising: one or more memories comprising instructions stored thereon (Prendki, ¶¶ [0235]-[00236] with 406/408/410 in FIG. 4: computer system 400 includes one or more units of memory 406 for electronically digitally storing data and instructions; computer system 400 further includes non-volatile memory such as read only memory (ROM) 408 or another static storage device; a unit of persistent storage 410 may include various forms of non-volatile RAM (NVRAM), such as FLASH memory, or solid-state storage, magnetic disk, or optical disk such as CD-ROM or DVDROM; storage 410 is an example of a non-transitory computer-readable medium that may be used to store instructions and data); and one or more processors configured to execute the instructions and perform operations described above (Prendki, ¶¶ [0234]-[0236] with 404 in FIG. 4: at least one hardware processor 404 for processing information and instructions; one or more units of memory 406, such as a main memory, for electronically digitally storing data and instructions to be executed by processor 404; instructions and data which when executed by the processor 404 cause performing computer-implemented methods to execute the techniques herein).
Prendki further discloses a non-transitory computer readable storage medium stored instructions thereon (Prendki ¶¶ [0235]-[00236] with 406/408/410 in FIG. 4: computer system 400 includes one or more units of memory 406 for electronically digitally storing data and instructions; computer system 400 further includes non-volatile memory such as read only memory (ROM) 408 or another static storage device; a unit of persistent storage 410 may include various forms of non-volatile RAM (NVRAM), such as FLASH memory, or solid-state storage, magnetic disk, or optical disk such as CD-ROM or DVDROM; storage 410 is an example of a non-transitory computer-readable medium that may be used to store instructions and data), the instructions, when executed by one or more processors, cause the one or more processors to perform operations described above (Prendki, ¶¶ [0234]-[0236] with 404 in FIG. 4: at least one hardware processor 404 for processing information and instructions; storage 410 is an example of a non-transitory computer-readable medium that may be used to store instructions and data which when executed by the processor 404 cause performing computer-implemented methods to execute the techniques herein).
Prendki fails to explicitly disclose dividing, based on the constraint, the labeling data of the target labeling task into automatic labeling data and manual labeling data.
Martin teaches a system and a method for labeling data (Martin, ¶ [0002]), wherein dividing, based on the constraint, the labeling data of the target labeling task into automatic labeling data and manual labeling data (Martin, ¶¶ [0025]-[0038]: provide mechanisms that can combine machine learning based data labeling with human specialist data labeling; as the machine learning component becomes more accurate, the labeling platform can automatically begin relying on the machine learning component more heavily, e.g., routing requests to human specialists when the machine learning component produces a low confidence result; a confidence-driven workflow (CDW) encapsulates a collection of labelers which are consulted in sequence, and their individual results are incorporated into an overall result, until a configured confidence threshold for an overall result is reached; e.g., the collection of labelers can include machine learning labelers, and human labelers, or combinations thereof; the constituent labelers are not directly linked, and the execution path is dynamically determined based, e.g., on confidence and cost constraint configuration; the order of consultation generally proceeds from least expensive labeler to most expensive labeler; in some cases, a given constituent labeler may be consulted more than once in the execution path; the labelers in a workflow may act as interfaces to labeler instances that are continuously monitored and scored based on the results produced by the labeler instances; the scores for the labeler instances behind a labeler can be used to dynamically determine a confidence range for the labeler and the confidence ranges for the labelers can be used in dynamic path determination; more particularly, the dynamic path determination mechanism can use the dynamically modeled confidence ranges for the labelers in the CDW to determine a priori confidence estimates for one or more paths and identify candidate paths that are predicted to meet cost and confidence constraints for a labeling request; the dynamic path determination mechanism can select a candidate path based, for example, on minimizing cost or other criteria; selection of a candidate path is based on expected cost and impact on overall result confidence; if the labeled result returned by a labeler does not match the cost and/or confidence expectations, the CDW can dynamically redetermine candidate paths; this redetermination can incorporate the accrued cost and overall result confidence estimate, the configured cost and confidence constraints, and the most recent available cost and confidence models of the CDW constituent labelers; after a step in the path, the labeling platform will have more information about the actual confidence and costs so far in the execution path and a redetermination of candidate paths can be performed help optimize the path from the current point in the execution path forward, whether or not the expectations for the current point in the path have been met; then, the (re)determination of candidate paths can occur for every step in the execution path, whether or not the expectations for that point have been met (e.g., until the overall all expectations for the execution path are met); dynamic path determination can account for the fact that the confidence in individual labelers may change; e.g., as more data is labeled, a machine learning labeler can be retrained, and the quality of the machine learning labeler goes up; consequently, the dynamically determined execution paths may increasingly terminate after a single consultation with the machine learning labeler, driving down the temporal and monetary costs of labeling by reducing reliance on human specialists; dynamically determining an execution path for a labeling request to label a data item, wherein dynamically determining the execution path comprises dynamically determining a bounded number of candidate paths through the set of labelers using dynamically calculated cost and confidence metrics for the labelers in the set of labelers to estimate a probability of each candidate path to satisfy a set of constraints on cost and a final result confidence; the next labeler consultation can be executed according to the selected path; receiving a labeled result from the selected labeler and determining a confidence estimate for the labeled result; the labeled result output of the selected labeler may be incorporated into an overall result; the overall result may be output as the final result for the CDW if the confidence estimate meets the constraint for the final result confidence; if the estimated confidence in the labeled result output by the labeler does not meet the target confidence threshold, the next labeler consultation can be redetermined; a new set of candidate paths can be redetermined using, for example, the accrued cost and result confidence estimate based on prior labeler consultations in the execution path and the confidence metrics for the labelers in the set of labelers to estimate a probability of each candidate path to satisfy a set of constraints on cost and final result confidence; continually monitoring and scoring a plurality of labeler instances to generate labeler instance scores for the plurality of labeler instances; scoring a labeler instance may include determining an accuracy of the labeler instance based on a correctness of a set labeled results produced by the labeler instance; updating the dynamically modeled confidence range for each labeler in the set of labelers; determining a set of labeler instance scores associated with a pool of labeler instances represented by the labeler and aggregating the set of labeler instance scores to generate the dynamically modeled confidence range for the labeler; a labeler may route labeling requests to labeler instances based on scores; ¶ [0057]-[0060] with FIG.2: Inputs that the labeler fails to label may be placed in an exception pipe 206; an element of input data may be considered a labeling request, which can comprise an element to be labeled or reference to the element to be labeled, such as an image or other data item to be labeled by the labeler; the labeling request may have associated flow control data, such as constraints on allowable confidence and cost, a list of labeler instances 203 acceptable to handle or not handle the request or other associated flow control information to control how the labeler 200 handles the request; ¶¶ [0095]-[0147] with FIGS. 7 and 8A-D: ML labelers, human labelers and other labelers can be combined into a confidence-driven workflow (CDW); a CDW can thus be considered a labeler that encapsulates a collection of other labelers, and more particularly, a collection of labelers of the same arity; the encapsulated labelers can be consulted in sequence and their individual results incorporated into an overall result until a configured threshold confidence target for an overall result is reached; a labeling request may be received as a workflow task with accompanying task information, such as a task type description and constraints; examples of constraints include, but are not limited to, cost constraints in one or more dimensions (e.g., time limit, monetary limit), target threshold confidence or other constraints; CDW 700 encapsulates ML labeler 712, blind judgement human labeler 714, and open judgement human labeler 716, though it should be appreciated that a CDW can encapsulate any number of labelers of various types; each labeler instance may have a labeler instance score determined, for example, by QMS 750, and that corresponds to the probability that the labeler instance will produce an accurate label for a given task; the QMS 750 can score how often the labeler instance was correct when it labeled images as "tumor" and score how often the labeler instance was correct when it labeled images as "no tumor"; the answer-specific scores can be used in determining confidence estimates for an actual result output by the labeler instance; human labeler instance scores and ML labeler instance scores can be tied to specific labeling task types in the scoring system; for a labeler instance that can produce labels for multiple task types, QMS 750 may determine scores on the different task types as skill scores to differentiate the labeler instance's performance across different labeling task types; a labeler instance may have an associated response time (temporal cost) that is an estimate of how long it will take that labeler instance to perform a task; a labeler instance may have an associated price (monetary cost) that is an estimate of the price for that labeler instance to perform a task; for a particular labeling task, the workflow orchestrator 710 may select from labelers suited for that type of task; a labeler may also have a labeler score; a labeler's score corresponds to the probability that a labeler will produce an accurate label for a given task; a workflow orchestrator 710 to dynamically determine an execution path for processing a labeling request ( question) to produce a final labeled result based on confidence and cost constraint configuration; workflow orchestrator 710 uses the task information for the workflow task, the labeler descriptions, possibly dynamic labeler characteristics such as cost, availability, and timeliness, and quality metrics from QMS 750 to dynamically determine a path through the constituent labelers 712, 714, 716 to produce a labeled output for the task which satisfies a configured target confidence threshold (a minimum confidence threshold); the path may also be selected to minimize costs in one or more dimensions; workflow orchestrator 710 uses the labeler scores and costs associated with the labelers to determine one or more viable paths through the labelers; once a viable path is determined, workflow orchestrator 710 may route the labeling request through that path until a result reaches a threshold confidence target or the path is exhausted; the path may be re-evaluated and changed at any time based on the actual results produced by each consultation; the order of consultation of constituent labelers may proceed from least expensive labeler to most expensive labeler, for example, in an attempt to reach the confidence threshold target with the least cost; if ML labeler 712 is the least expensive labeler, blind judgement human labeler 714 is more expensive because labeling requests to blind judgement human labeler 714 are routed to human specialists who have a higher monetary cost based on compensation for work performed, and human labeler 716 is the most expensive of the constituent labelers because labeling requests to human labeler 716 are routed to human specialists with higher expertise and compensation levels than the human specialists associated with blind judgement human labeler 714, then workflow orchestrator 710 may favor consulting ML labeler 712 first, then human labeler 714, then human labeler 716; in the example of FIG. 8B, the labeling request is routed to ML labeler 712, which returns an answer for which a confidence estimate 806 is determined; the labeling request is then routed to blind judgement human labeler 714, which produces an answer for which a confidence estimate 808 is determined; as the target confidence threshold has not been reached, the labeling request can be routed to blind judgement human labeler 714 twice because there are labeler instances (human specialists 727) remaining in the pool of human labeler 714 that have not yet been consulted for the labeling request; the confidence estimate 810 for the answer produced by the second consultation with human labeler 714 may incorporate confidence estimates 806, 808, however, exceeds the target confidence threshold; thus, that answer can be used for the final labeled results of CDW 700; QMS 750 provides quality metrics that may be used to dynamically determine the execution path and determine if a label result received from a constituent labeler ( or agreed to by multiple labelers) exceeds the configured target confidence threshold; QMS 750 can continuously monitor and score labeler instances over time to generate, maintain, and improve confidence estimates; QMS 750 can identify the set of constraints on labelers required to achieve a desired confidence threshold for an overall result (which may be composed of results from multiple labelers); various methods of determining the complex granular confidence estimate may be used; e.g., approaches may be employed that take into account the actual value of the answer (e.g., precision versus recall considerations, which into account differences in false positives versus false negatives and different likelihoods of providing one wrong answer versus a different wrong answer when the true answer is x versus if it is y); logistic regression methods or other estimators may be applied to determine complex granular confidence estimates rather than closed form probability equations; an ML estimation model may be used for combining into a single probability value; logistic regression models can be trained for combining confidence estimates into a single probability value).
Prendki and Martin are analogous art because they are from the same field of endeavor, a system and a method for labeling data . Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of Martin to Prendki. Motivation for doing so would balance the predicted benefit against the costs associated with each labeler.
Claim 2
Prendki in view of Martin discloses all the elements as stated in Claim 1 and further discloses wherein determining the labeling procedure based on the data labeling request and at least one selected from a group consisting of the task type, the labeling quality, and the labeling metric comprises: determining, based on a task type of the target labeling task, whether an existing labeling procedure corresponding to the target labeling task is present; if the existing labeling procedure is present, determining whether the existing labeling procedure satisfies a first quality evaluation index; and if the existing labeling procedure is not present: obtaining a set of automatic labeling procedures based on the task type of the target labeling task; and determining, based on the set of automatic labeling procedures, the labeling procedure (Prendki, ¶ [0050]: efficiently compare the capabilities, availability, performance, and costs of various labeling providers; a computer-implemented labeling marketplace, through which a suitable labeling partner (i.e., labeling provider partnered with the computer-implemented labeling marketplace) will be automatically selected, based both on historical and real-time metrics, to handle the task requested by a customer; provide a short-list of labeling partners that are similarly determined to be suitable for handling the task; such selection processes may involve eliminating from the selection pool labeling partners that are not capable of handling the requested task; determining real-time availability of the labeling partners, and evaluating the labeling partners' performance history, costs, and various other factors that may be relevant to delivering on the expectations of the customers; incorporate a quality control process, in which the tasks completed by the labeling partners may be evaluated; ¶¶ [0056]-[0057], [0059], [0065], [0067],and [0110] with FIG. 3: recommendation computer system 306 is programmed using the real-time evaluation instructions 308 to query the provider metrics database 310 and to select, automatically and based upon programmed recommendation or evaluation algorithms, one or more of the label service providers 314, 316, 318 that is optimal to perform one or more labeling tasks on the user dataset; provides labeling services to customers by finding a labeling partner that is suitable to handle the tasks requested by the customers; a customer may rely on the system to automatically select the labeling partner that is most likely to provide the highest labeling quality in a timely manner, or alternatively, the system may provide the customer with a short-list of labeling partners determined to be suitable for handling the requested task; many labeling tasks require specialized knowledge in particular technologies or languages (e.g., computer vision, translation) or access to a specific kind of tools (e.g., software or hardware for Light Detection and Ranging, or LiDAR); some of these specialized or smaller labeling providers are able to provide superior labeling accuracy and faster tum-around time for certain types of tasks than their larger counterparts because they employ annotators that are more capable, knowledgeable, and skilled in handling those types of tasks; allowing customers to rely on a computer-implemented labeling marketplace to find the labeling provider that is most likely to produce the highest labeling quality and in a timely-manner, regardless of the size of the labeling providers and for whatever types of labeling tasks; labeling partners unable to perform the labeling task because of an inability to perform tasks with the data type and task type requested may be eliminated from consideration for the labeling task; the minimum criteria may be determined during one of the initial phases of the system's processes so any labeling partners that are not capable of handling the requested tasks could be eliminated from the selection pool; ¶¶ [0146]-[0154]: utilize a multi-objective optimization process to find a labeling partner that is most likely to produce the highest quality of work; examples of the factors considered during the optimization process include, but are not limited to, the following: a) turn-around time (e.g., on average how long does the labeling partner need to complete and return a task); b) labeling quality (e.g., how accurate is the labeled data); c) average cost (e.g., the average cost charged by the labeling partner for various types of tasks); d) revenue (e.g., how much did the labeling partner earn for a given period of time); e) number of, and type of, tasks previously offered to the labeling partner; f) number of labeling partners within the computer-implemented labeling marketplace; and g) customer's preferences/criteria; the optimization process may be programmed using any of: the Dantzig simplex algorithm, extensions or variants thereof; combinatorial algorithms; or quantum optimization algorithms; ¶ [0168]: identify a labeling partner that is most suitable for handling the customer's task by considering, among others, real-time availability of the annotators employed by the various labeling partners; allows the system to track each of their annotator's computer activity, or to provide some other indicators to verify whether the annotators are available to receive a labeling task; ¶¶ [0169]-[0173] with FIG. 5F: to provide customers with consistent and efficient turn-around time for the delivery of the requested tasks, the system may place certain time constraints on the labeling partners; the system may provide the labeling partner a limited amount of time to complete any task accepted by the labeling partner, which time period may be determined based on the historical data of the labeling partner or other labeling partners (e.g., average amount of time needed to complete a task of a similar type); the customer may be shown a ranked list of labeling partners and be requested to select its top three choices; the first selected labeling partner may be offered the labeling task; if the first selected labeling partner fails to respond in time, the system may automatically route the request to the labeling partner that was the customer's second top choice; if all three of the customer's initial selections decline the request or fail to timely respond, the system may notify the customer of this failed status and may prompt the customer to continue to select additional labeling partners or to make adjustments to its labeling methodology until a suitable labeling partner accepts the offer for the labeling task; FIG. 5F illustrates a computer display device that is displaying an example of a graphical user interface with which a customer can select one or more labeling partners and view recommendations of labeling partners; ¶¶ [0176]-[0183] with FIG. 6B: select a labeling partner that is suitable for the customer's requested task in a sequential manner, such that the task is offered to only one labeling partner at a time; the system may re-evaluate all of the labeling partners within the selection pool, e.g., in situations where enough time has passed, and real-time metrics associated with the labeling partners have substantially changed; select a labeling partner that is suitable for the customer's requested task by employing a self-serve process, bidding process, or reverse-bidding process; if, e.g., a customer prioritizes cost as an important criteria for the task, the task may be offered to several labeling partners at the same time; in such a case, the labeling partner that provides the lowest cost for the task may be selected; alternatively, the system may provide to the customer a list of labeling providers determined to be suitable for the task, along with the costs associated with each of the providers; the system (which has historical data for price, time and accuracy from previous jobs, on a use case-by-use case basis), is programmed to compute the z-score of each valid labeling partner for price, time and accuracy; each of the z-scores is weighted with the number of points allocated to the feature (for price and time, a -1 factor is added, since lower price or time is better); the system is programmed to compute the sum and to return a ranked list of recommendations based on the sum values; FIG. 6B illustrates an example of a graphical user interface that can be programmed as presentation output of an ordered set of recommended labeling partners; provide task-level feedback that provides information about why a specific task was not offered to a labeling partner (e.g., insufficient availability of annotators, lack of tooling, etc.) or aggregated feedback that provides information about why specific types of tasks, or tasks in general, are not offered to the labeling partner (e.g., statistical insights about subpar turn-around time or labeling accuracy); ¶¶ [0208]-[0209]: select a labeling partner to handle the labeling at an experiment-level (i.e., the task in its entirety), which may involve one or more loops of labeling; this will provide consistency in the labeled data and minimize the inefficiencies that come from transitioning the task between several labeling partners; select several labeling partners to handle the task, each labeling partner being assigned to a particular loop of the task; e.g., if the system determines that a task can be partitioned into several loops and some of the loops are better suitable for one labeling partner while other loops are better suitable for another labeling partner, the system may assign the task to both labeling partners; if the system determines that a labeling task includes an amount of data that is too large for any one of the labeling partners in the selection pool, the system may assign various portions of the task to multiple labeling partners; the system may also assign a labeling task to several labeling partners if it determines that diversity among annotators may produce higher labeling quality; ¶¶ [0218]-[0221] with FIG. 1: to provide labeling services to a customer, the labeling partner needs to fulfill the following expectations: (a) possess a task force capable/trained to manage the particular type of task; (b) possess the proper annotation tool for the said task; be compliant to the level required by the user; (c) be compliant to the level required by the user; and (d) have at least one annotator available immediately; at step 102, the system applies the minimum criteria to each of the pre-selected labeling partners within the selection pool and eliminates any labeling partners that do not meet the minimum criteria, i.e., that are not capable of handling the task; at steps 103 and 104, the system applies an optimization process that evaluates real-time and historical metrics of all the labeling partners remaining in the selection pool; at step 103, check the real-time availability of the labeling partners, which may include checking real-time availability of the annotators employed by each of the labeling partners; at step 104, evaluates the historical metrics along with the real-time metrics; examples of the historical metrics evaluated by the system may include each of the labeling partners' average labeling accuracy, average turn-around time, and average pricing, both at an overall level and task-specific level (e.g., by grouping together similar types of tasks); based on the real-time and historical metrics, identify a labeling partner that is the most likely to produce the highest quality of work and in a timely-manner; at step 109, if a labeling partner either refuses the task or does not reply within a set amount of time, remove the labeling partner from the selection pool and identifies another labeling partner that is suitable for handling the task; identify the subsequent labeling partner by re-evaluating the real-time and historical metrics of the labeling partners left in the selection pool; alternatively, identify the subsequent labeling partner based on the previous evaluation, e.g., in situations where insignificant amount of time has passed since the previous evaluation; ¶ [0229] with FIGS. 1-2: the steps of identifying/selecting an optimal labeling partner (i.e., steps 102-108 within brackets 260) may be skipped for loops n> 1, if the system determines that the same labeling partner should be used throughout the experiment; alternatively, the steps within brackets 260 may be kept in place if the system determines that different labeling partners should be used within the same experiment) (Martin, ¶¶ [0085]-[0094] with FIG. 6: the ML labeler may be implemented as a wrapper for an ML model on an ML platform 650 running locally or on a remote ML platform system; the ML labeler configuration can specify an ML algorithm to use and, based on the ML algorithm specified, labeling platform 104 configures the labeler with the code to connect to the appropriate ML platform 650 to train and use the specified ML algorithm; ML labeler 600 includes training component 615 executable to train an ML algorithm; the training component 615 includes an experiment coordinator 616 that interfaces with ML platform 650 to train multiple challenger models ( e.g., using various hyperparameters or other mechanisms for training multiple candidate models known or developed in the art) and a challenger model evaluator 618 that evaluates candidate ML models against each other and the current active model to determine which should be the current active model for inferring answers to labeling requests; the output is a champion ML model that represents the best model currently producible given the available training data; the training component 615 thus determines the ML model to use as the current active model for inferring answers to labeling requests; training triggers may be based on, for example, an amount of training data received by the labeler, quality metrics received by the labeler, elapsed time, or other criteria).
Claim 3
Prendki in view of Martin discloses all the elements as stated in Claim 2 and further discloses wherein determining, based on the set of automatic labeling procedures, the labeling procedure comprises: computing a labeling quality and a labeling metric of each automatic labeling procedure in the set of automatic labeling procedures; determining a first automatic labeling procedure based on the labeling quality and the labeling metric; determining whether the first automatic labeling procedure satisfies the first quality evaluation index; if the first automatic labeling procedure satisfies the first quality evaluation index, setting the first automatic labeling procedure as the labeling procedure (Prendki, ¶¶ [0176]-[0183] with FIG. 6B: select a labeling partner that is suitable for the customer's requested task in a sequential manner, such that the task is offered to only one labeling partner at a time; the system may re-evaluate all of the labeling partners within the selection pool, e.g., in situations where enough time has passed, and real-time metrics associated with the labeling partners have substantially changed; select a labeling partner that is suitable for the customer's requested task by employing a self-serve process, bidding process, or reverse-bidding process; if, e.g., a customer prioritizes cost as an important criteria for the task, the task may be offered to several labeling partners at the same time; in such a case, the labeling partner that provides the lowest cost for the task may be selected; alternatively, the system may provide to the customer a list of labeling providers determined to be suitable for the task, along with the costs associated with each of the providers; the system (which has historical data for price, time and accuracy from previous jobs, on a use case-by-use case basis), is programmed to compute the z-score of each valid labeling partner for price, time and accuracy; each of the z-scores is weighted with the number of points allocated to the feature (for price and time, a -1 factor is added, since lower price or time is better); the system is programmed to compute the sum and to return a ranked list of recommendations based on the sum values; FIG. 6B illustrates an example of a graphical user interface that can be programmed as presentation output of an ordered set of recommended labeling partners; provide task-level feedback that provides information about why a specific task was not offered to a labeling partner (e.g., insufficient availability of annotators, lack of tooling, etc.) or aggregated feedback that provides information about why specific types of tasks, or tasks in general, are not offered to the labeling partner (e.g., statistical insights about subpar turn-around time or labeling accuracy); ¶¶ [0218]-[0221] with FIG. 1: to provide labeling services to a customer, the labeling partner needs to fulfill the following expectations: (a) possess a task force capable/trained to manage the particular type of task; (b) possess the proper annotation tool for the said task; be compliant to the level required by the user; (c) be compliant to the level required by the user; and (d) have at least one annotator available immediately; at step 102, the system applies the minimum criteria to each of the pre-selected labeling partners within the selection pool and eliminates any labeling partners that do not meet the minimum criteria, i.e., that are not capable of handling the task; at steps 103 and 104, the system applies an optimization process that evaluates real-time and historical metrics of all the labeling partners remaining in the selection pool; at step 103, check the real-time availability of the labeling partners, which may include checking real-time availability of the annotators employed by each of the labeling partners; at step 104, evaluates the historical metrics along with the real-time metrics; examples of the historical metrics evaluated by the system may include each of the labeling partners' average labeling accuracy, average turn-around time, and average pricing, both at an overall level and task-specific level (e.g., by grouping together similar types of tasks); based on the real-time and historical metrics, identify a labeling partner that is the most likely to produce the highest quality of work and in a timely-manner; at step 109, if a labeling partner either refuses the task or does not reply within a set amount of time, remove the labeling partner from the selection pool and identifies another labeling partner that is suitable for handling the task; identify the subsequent labeling partner by re-evaluating the real-time and historical metrics of the labeling partners left in the selection pool; alternatively, identify the subsequent labeling partner based on the previous evaluation, e.g., in situations where insignificant amount of time has passed since the previous evaluation) (Martin, ¶¶ [0095]-[0122] with FIGS. 7 and 8A-D: ML labelers, human labelers and other labelers can be combined into a confidence-driven workflow (CDW); a CDW can thus be considered a labeler that encapsulates a collection of other labelers, and more particularly, a collection of labelers of the same arity; the encapsulated labelers can be consulted in sequence and their individual results incorporated into an overall result until a configured threshold confidence target for an overall result is reached; a labeling request may be received as a workflow task with accompanying task information, such as a task type description and constraints; examples of constraints include, but are not limited to, cost constraints in one or more dimensions (e.g., time limit, monetary limit), target threshold confidence or other constraints; CDW 700 encapsulates ML labeler 712, blind judgement human labeler 714, and open judgement human labeler 716, though it should be appreciated that a CDW can encapsulate any number of labelers of various types; each labeler instance may have a labeler instance score determined, for example, by QMS 750, and that corresponds to the probability that the labeler instance will produce an accurate label for a given task; the QMS 750 can score how often the labeler instance was correct when it labeled images as "tumor" and score how often the labeler instance was correct when it labeled images as "no tumor"; the answer-specific scores can be used in determining confidence estimates for an actual result output by the labeler instance; human labeler instance scores and ML labeler instance scores can be tied to specific labeling task types in the scoring system; for a labeler instance that can produce labels for multiple task types, QMS 750 may determine scores on the different task types as skill scores to differentiate the labeler instance's performance across different labeling task types; a labeler instance may have an associated response time (temporal cost) that is an estimate of how long it will take that labeler instance to perform a task; a labeler instance may have an associated price (monetary cost) that is an estimate of the price for that labeler instance to perform a task; for a particular labeling task, the workflow orchestrator 710 may select from labelers suited for that type of task; a labeler may also have a labeler score; a labeler's score corresponds to the probability that a labeler will produce an accurate label for a given task; a workflow orchestrator 710 to dynamically determine an execution path for processing a labeling request ( question) to produce a final labeled result based on confidence and cost constraint configuration; workflow orchestrator 710 uses the task information for the workflow task, the labeler descriptions, possibly dynamic labeler characteristics such as cost, availability, and timeliness, and quality metrics from QMS 750 to dynamically determine a path through the constituent labelers 712, 714, 716 to produce a labeled output for the task which satisfies a configured target confidence threshold (a minimum confidence threshold); the path may also be selected to minimize costs in one or more dimensions; workflow orchestrator 710 uses the labeler scores and costs associated with the labelers to determine one or more viable paths through the labelers; once a viable path is determined, workflow orchestrator 710 may route the labeling request through that path until a result reaches a threshold confidence target or the path is exhausted; the path may be re-evaluated and changed at any time based on the actual results produced by each consultation; the order of consultation of constituent labelers may proceed from least expensive labeler to most expensive labeler, for example, in an attempt to reach the confidence threshold target with the least cost; if ML labeler 712 is the least expensive labeler, blind judgement human labeler 714 is more expensive because labeling requests to blind judgement human labeler 714 are routed to human specialists who have a higher monetary cost based on compensation for work performed, and human labeler 716 is the most expensive of the constituent labelers because labeling requests to human labeler 716 are routed to human specialists with higher expertise and compensation levels than the human specialists associated with blind judgement human labeler 714, then workflow orchestrator 710 may favor consulting ML labeler 712 first, then human labeler 714, then human labeler 716; in the example of FIG. 8B, the labeling request is routed to ML labeler 712, which returns an answer for which a confidence estimate 806 is determined; the labeling request is then routed to blind judgement human labeler 714, which produces an answer for which a confidence estimate 808 is determined; as the target confidence threshold has not been reached, the labeling request can be routed to blind judgement human labeler 714 twice because there are labeler instances (human specialists 727) remaining in the pool of human labeler 714 that have not yet been consulted for the labeling request; the confidence estimate 810 for the answer produced by the second consultation with human labeler 714 may incorporate confidence estimates 806, 808, however, exceeds the target confidence threshold; thus, that answer can be used for the final labeled results of CDW 700; ¶¶ [0181]-[0215] with FIGS. 11A-B: route tasks to labelers based on any number of constraints that match the labelers' descriptions; specific sequencing of labelers can be specified to achieve predefined workflows; at step 1102, workflow orchestrator 710 applies criteria to filter out labelers from consideration; if there are cost constraints (temporal, monetary or other cost) and/or confidence constraints, workflow orchestrator 710 performs a path search to find the optimal set of labelers (step 1114); each path endpoint can represent a specific path of labelers (that is, a sequence of labelers consulted) and QMS 750 estimates the a priori endpoint confidence for the path (step 1120) before the path is executed; if the path meets the task constraints based on the estimated total path monetary cost, total path time, and end-point confidence, the path can be added to a set of viable paths (step 1120); otherwise the path can be discarded (step 1122); if at least one viable path is found, workflow orchestrator can select a path from the one or more viable paths (step 1128); if the selected viable path includes only one labeler (one node) (as determined at step 1129), the minimum confidence required for the labeler in order to maintain viability of the planned path may be determined (step 1130); at step 1132, workflow orchestrator 710 sends the task to that labeler with a confidence constraint, the confidence constraint including the required minimum confidence determined by QMS 750; at any point, if all remaining constraints and optimizations are satisfied by multiple labelers, then a random selection is made from these labelers; workflow orchestrator 710 selects the first labeler in a selected path (step 1134) and sends the task to that labeler with or without a confidence constraint (step 1136); at step 1138, workflow orchestrator 710 receives an output of the labeler; workflow orchestrator 710 determines the confidence estimate for the label returned by the labeler (step 1142); the confidence estimate returned for the label by QMS 750 meets the confidence threshold target, as determined at step 1144, then a stopping condition has been reached, and workflow orchestrator 710 can return the labeled result, including the confidence determined by QMS 750 for the label (step 1146)).
Claim 4
Prendki in view of Martin discloses all the elements as stated in Claim 1 and further discloses wherein determining the target labeling task based on the one or more labeling index values and the one or more labeling tasks comprises: computing one or more first numbers corresponding to one or more labeling tasks by at least: for each labeling task in the one or more labeling tasks, computing a first number of a rest of the labeling tasks dominated by the each labeling task based on a corresponding labeling index value of the each labeling task, wherein the rest of the labeling tasks are one or more remaining labeling tasks after selecting a labeling task from the one or more labeling tasks; sorting a priority of one or more labeling tasks based on the one or more first numbers; and determining a labeling task with the highest first number in the one or more first numbers as the target labeling task (Prendki, ¶¶ [0060], [0174], and [0226]-[0229] with FIGS. 1-2: at step 201, a customer uploads to the system labeling instructions, data to be labeled, and a model to be trained using the data; at step 202, the customer decides a labeling approach (e.g., ML-based approach, manual approach, or self-labeling approach); at step 203, the curation process begins, and at step 204, the customer's model and data are sent to the data curation engine; the system's data preparation and data curation services employ an Active Learning approach where portions of the data are prioritized so the portions that are most effective in training the model are labeled first; the curation engine analyzes the data and identifies a first batch of data that it believes will be effective in training the data; the first batch of data is provided to a labeling partner in a similar fashion as the processes described in FIG. 1; once the labeling partner completes labeling the first batch of data, the labeled data is used to train the customer's model, as described in step 206; then, the model's performance is evaluated; at steps 207 and 208, if the system determines that the model's performance is acceptable, the system deems the model as sufficiently trained and provides the trained model and labeled data back to the customer (i.e., the experiment ends); the experiment may also end if the system determines that the customer's budget has ran out or if the model has reached its optimal performance; at step 205, if the system determines that the model's performance is unacceptable, the next loop of the experiment begins; going back to step 204, the data curation engine identifies the next batch of data and provides this batch to the labeling partner along with all of the previous batches of data; once the labeling partner completes labeling a batch of the data, the experiment continues until the model's performance reaches an acceptable level, the budget is reached, or optimal performance has been achieved because the remaining data is determined as redundant or irrelevant for training the model; ¶ [0214]: use a Ranking-based Active Learning approach; with Ranking-based Active Learning, the system is programmed to re-rank data; rather than predicting which record should be labeled in the context of a specific loop, the system is programmed to continuously re-rank records in the dataset based on qualitative aspects of the records (e.g., how effective the dataset is in training the model), which allows the system to select data not only for loop n, but the subsequent ones most effective in training the model as well).
Claim 5
Prendki in view of Martin discloses all the elements as stated in Claim 1 and further discloses wherein dividing, based on the constraint, the labeling data of the target labeling task into automatic labeling data and manual labeling data comprises: deciding whether each data item of the one or more data items in the labeling data satisfies each constraint condition of the one or more constraint conditions of the constraint; assigning one or more data items of the labeling data satisfying the constraint as the automatic labeling data; assigning one or more data items of the labeling data not satisfying the constraint as the manual labeling data (Martin, ¶¶ [0095]-[0122] with FIGS. 7 and 8A-D: ML labelers, human labelers and other labelers can be combined into a confidence-driven workflow (CDW); a CDW can thus be considered a labeler that encapsulates a collection of other labelers, and more particularly, a collection of labelers of the same arity; the encapsulated labelers can be consulted in sequence and their individual results incorporated into an overall result until a configured threshold confidence target for an overall result is reached; a labeling request may be received as a workflow task with accompanying task information, such as a task type description and constraints; examples of constraints include, but are not limited to, cost constraints in one or more dimensions (e.g., time limit, monetary limit), target threshold confidence or other constraints; CDW 700 encapsulates ML labeler 712, blind judgement human labeler 714, and open judgement human labeler 716, though it should be appreciated that a CDW can encapsulate any number of labelers of various types; each labeler instance may have a labeler instance score determined, for example, by QMS 750, and that corresponds to the probability that the labeler instance will produce an accurate label for a given task; the QMS 750 can score how often the labeler instance was correct when it labeled images as "tumor" and score how often the labeler instance was correct when it labeled images as "no tumor"; the answer-specific scores can be used in determining confidence estimates for an actual result output by the labeler instance; human labeler instance scores and ML labeler instance scores can be tied to specific labeling task types in the scoring system; for a labeler instance that can produce labels for multiple task types, QMS 750 may determine scores on the different task types as skill scores to differentiate the labeler instance's performance across different labeling task types; a labeler instance may have an associated response time (temporal cost) that is an estimate of how long it will take that labeler instance to perform a task; a labeler instance may have an associated price (monetary cost) that is an estimate of the price for that labeler instance to perform a task; for a particular labeling task, the workflow orchestrator 710 may select from labelers suited for that type of task; a labeler may also have a labeler score; a labeler's score corresponds to the probability that a labeler will produce an accurate label for a given task; a workflow orchestrator 710 to dynamically determine an execution path for processing a labeling request ( question) to produce a final labeled result based on confidence and cost constraint configuration; workflow orchestrator 710 uses the task information for the workflow task, the labeler descriptions, possibly dynamic labeler characteristics such as cost, availability, and timeliness, and quality metrics from QMS 750 to dynamically determine a path through the constituent labelers 712, 714, 716 to produce a labeled output for the task which satisfies a configured target confidence threshold (a minimum confidence threshold); the path may also be selected to minimize costs in one or more dimensions; workflow orchestrator 710 uses the labeler scores and costs associated with the labelers to determine one or more viable paths through the labelers; once a viable path is determined, workflow orchestrator 710 may route the labeling request through that path until a result reaches a threshold confidence target or the path is exhausted; the path may be re-evaluated and changed at any time based on the actual results produced by each consultation; the order of consultation of constituent labelers may proceed from least expensive labeler to most expensive labeler, for example, in an attempt to reach the confidence threshold target with the least cost; if ML labeler 712 is the least expensive labeler, blind judgement human labeler 714 is more expensive because labeling requests to blind judgement human labeler 714 are routed to human specialists who have a higher monetary cost based on compensation for work performed, and human labeler 716 is the most expensive of the constituent labelers because labeling requests to human labeler 716 are routed to human specialists with higher expertise and compensation levels than the human specialists associated with blind judgement human labeler 714, then workflow orchestrator 710 may favor consulting ML labeler 712 first, then human labeler 714, then human labeler 716; in the example of FIG. 8B, the labeling request is routed to ML labeler 712, which returns an answer for which a confidence estimate 806 is determined; the labeling request is then routed to blind judgement human labeler 714, which produces an answer for which a confidence estimate 808 is determined; as the target confidence threshold has not been reached, the labeling request can be routed to blind judgement human labeler 714 twice because there are labeler instances (human specialists 727) remaining in the pool of human labeler 714 that have not yet been consulted for the labeling request; the confidence estimate 810 for the answer produced by the second consultation with human labeler 714 may incorporate confidence estimates 806, 808, however, exceeds the target confidence threshold; thus, that answer can be used for the final labeled results of CDW 700).
Claim 6
Prendki in view of Martin discloses all the elements as stated in Claim 5 and further discloses wherein assigning one or more data items of the labeling data not satisfying the constraint as the manual labeling data comprises: determining a target pre-filter procedure based on a task type of the target labeling task; determining, respectively, whether each data item of the one or more data items in the manual labeling data satisfies a target pre-filter procedure threshold based on the target pre-filter procedure; sending one or more data items satisfying the target pre-filter procedure threshold to a manual labeling data item queue; discarding one or more data items not satisfying the target pre-filter procedure threshold (Prendki, ¶¶ [0065], [0067], and [0110] many labeling tasks require specialized knowledge in particular technologies or languages (e.g., computer vision, translation) or access to a specific kind of tools (e.g., software or hardware for Light Detection and Ranging, or LiDAR); some of these specialized or smaller labeling providers are able to provide superior labeling accuracy and faster tum-around time for certain types of tasks than their larger counterparts because they employ annotators that are more capable, knowledgeable, and skilled in handling those types of tasks; allowing customers to rely on a computer-implemented labeling marketplace to find the labeling provider that is most likely to produce the highest labeling quality and in a timely-manner, regardless of the size of the labeling providers and for whatever types of labeling tasks; determine the minimum criteria required for handling the labeling tasks requested by the customers; labeling partners unable to perform the labeling task because of an inability to perform tasks with the data type and task type requested may be eliminated from consideration for the labeling task; the minimum criteria may be determined during one of the initial phases of the system's processes so any labeling partners that are not capable of handling the requested tasks could be eliminated from the selection pool; ¶¶ [0191]-[0200] with FIG. 8: Hybrid Labeling Marketplace: find awesome labeling partners that are available immediately; process flow: 1. tell us about your project, expectations, and time and budgetary constraints; 2. choose between the recommended auto-labeling and manual labeling options; Mix-and-Match; 3. you get a precise estimate of the cost, time and accuracy; once the task launches, your data is immediately being labeled; 4. use the human-in-the-loop module to find mislabeled data and send it for revision either to the same partner or a different partner; human-in-the-loop module: visualize, audit, and fix your labels and measure labeling quality reliably; 1. upload your existing labels or select a labeling task after the labels are sent back by your labeling partner; 2. view label anomalies or query data using your own validation criteria; no need to ever review an entire dataset anymore; 3. fix faulty labels yourself with an annotation tool or send them back to either the same or a different labeling partner; 4. as you perfect your labels, they are automatically versioned, so that you can revert to an older version if needed; ¶ [0209]: if the system determines that a labeling task includes an amount of data that is too large for any one of the labeling partners in the selection pool, the system may assign various portions of the task to multiple labeling partners; the system may also assign a labeling task to several labeling partners if it determines that diversity among annotators may produce higher labeling quality) (Martin, ¶¶ [0055]-[0068] with FIGS. 2-3: Labelers (including labelers of different types) can be composed together into directed graphs as needed, such that each individual labeler solves a portion of an overall classification problem, and the results are aggregated together to form the overall labeled output; Labeling platform 104 may include multiple types of labelers and multiple labelers of each type; translation by conditioning layer 304 may be required because the data domain external to the kernel core logic 302 may be different than the kernel's data domain; the external data domain may be use-case specific and technology agnostic, while the kernel's data domain may be technology-specific and use case agnostic; the conditioning layer 304 may also perform validation on inbound data; e.g., for one use case, a solid black image may be valid for training/inferring, while for other use cases, it may not; if it is not, the conditioning layer 304 may, for example, include a filter to remove solid black images; alternatively, it might reject such input and issue an exception output; ¶¶ [0095]-[0122] with FIGS. 7 and 8A-D: ML labelers, human labelers and other labelers can be combined into a confidence-driven workflow (CDW); a CDW can thus be considered a labeler that encapsulates a collection of other labelers, and more particularly, a collection of labelers of the same arity; the encapsulated labelers can be consulted in sequence and their individual results incorporated into an overall result until a configured threshold confidence target for an overall result is reached; a labeling request may be received as a workflow task with accompanying task information, such as a task type description and constraints; examples of constraints include, but are not limited to, cost constraints in one or more dimensions (e.g., time limit, monetary limit), target threshold confidence or other constraints; CDW 700 encapsulates ML labeler 712, blind judgement human labeler 714, and open judgement human labeler 716, though it should be appreciated that a CDW can encapsulate any number of labelers of various types; each labeler instance may have a labeler instance score determined, for example, by QMS 750, and that corresponds to the probability that the labeler instance will produce an accurate label for a given task; the QMS 750 can score how often the labeler instance was correct when it labeled images as "tumor" and score how often the labeler instance was correct when it labeled images as "no tumor"; the answer-specific scores can be used in determining confidence estimates for an actual result output by the labeler instance; human labeler instance scores and ML labeler instance scores can be tied to specific labeling task types in the scoring system; for a labeler instance that can produce labels for multiple task types, QMS 750 may determine scores on the different task types as skill scores to differentiate the labeler instance's performance across different labeling task types; a labeler instance may have an associated response time (temporal cost) that is an estimate of how long it will take that labeler instance to perform a task; a labeler instance may have an associated price (monetary cost) that is an estimate of the price for that labeler instance to perform a task; for a particular labeling task, the workflow orchestrator 710 may select from labelers suited for that type of task; a labeler may also have a labeler score; a labeler's score corresponds to the probability that a labeler will produce an accurate label for a given task; a workflow orchestrator 710 to dynamically determine an execution path for processing a labeling request ( question) to produce a final labeled result based on confidence and cost constraint configuration; workflow orchestrator 710 uses the task information for the workflow task, the labeler descriptions, possibly dynamic labeler characteristics such as cost, availability, and timeliness, and quality metrics from QMS 750 to dynamically determine a path through the constituent labelers 712, 714, 716 to produce a labeled output for the task which satisfies a configured target confidence threshold (a minimum confidence threshold); the path may also be selected to minimize costs in one or more dimensions; workflow orchestrator 710 uses the labeler scores and costs associated with the labelers to determine one or more viable paths through the labelers; once a viable path is determined, workflow orchestrator 710 may route the labeling request through that path until a result reaches a threshold confidence target or the path is exhausted; the path may be re-evaluated and changed at any time based on the actual results produced by each consultation; the order of consultation of constituent labelers may proceed from least expensive labeler to most expensive labeler, for example, in an attempt to reach the confidence threshold target with the least cost; if ML labeler 712 is the least expensive labeler, blind judgement human labeler 714 is more expensive because labeling requests to blind judgement human labeler 714 are routed to human specialists who have a higher monetary cost based on compensation for work performed, and human labeler 716 is the most expensive of the constituent labelers because labeling requests to human labeler 716 are routed to human specialists with higher expertise and compensation levels than the human specialists associated with blind judgement human labeler 714, then workflow orchestrator 710 may favor consulting ML labeler 712 first, then human labeler 714, then human labeler 716; in the example of FIG. 8B, the labeling request is routed to ML labeler 712, which returns an answer for which a confidence estimate 806 is determined; the labeling request is then routed to blind judgement human labeler 714, which produces an answer for which a confidence estimate 808 is determined; as the target confidence threshold has not been reached, the labeling request can be routed to blind judgement human labeler 714 twice because there are labeler instances (human specialists 727) remaining in the pool of human labeler 714 that have not yet been consulted for the labeling request; the confidence estimate 810 for the answer produced by the second consultation with human labeler 714 may incorporate confidence estimates 806, 808, however, exceeds the target confidence threshold; thus, that answer can be used for the final labeled results of CDW 700; ¶¶ [0181]-[0215] with FIGS. 11A-B: route tasks to labelers based on any number of constraints that match the labelers' descriptions; specific sequencing of labelers can be specified to achieve predefined workflows; at step 1102, workflow orchestrator 710 applies criteria to filter out labelers from consideration; if there are cost constraints (temporal, monetary or other cost) and/or confidence constraints, workflow orchestrator 710 performs a path search to find the optimal set of labelers (step 1114); each path endpoint can represent a specific path of labelers (that is, a sequence of labelers consulted) and QMS 750 estimates the a priori endpoint confidence for the path (step 1120) before the path is executed; if the path meets the task constraints based on the estimated total path monetary cost, total path time, and end-point confidence, the path can be added to a set of viable paths (step 1120); otherwise the path can be discarded (step 1122); if at least one viable path is found, workflow orchestrator can select a path from the one or more viable paths (step 1128); if the selected viable path includes only one labeler (one node) (as determined at step 1129), the minimum confidence required for the labeler in order to maintain viability of the planned path may be determined (step 1130); at step 1132, workflow orchestrator 710 sends the task to that labeler with a confidence constraint, the confidence constraint including the required minimum confidence determined by QMS 750; at any point, if all remaining constraints and optimizations are satisfied by multiple labelers, then a random selection is made from these labelers; workflow orchestrator 710 selects the first labeler in a selected path (step 1134) and sends the task to that labeler with or without a confidence constraint (step 1136); at step 1138, workflow orchestrator 710 receives an output of the labeler; workflow orchestrator 710 determines the confidence estimate for the label returned by the labeler (step 1142); the confidence estimate returned for the label by QMS 750 meets the confidence threshold target, as determined at step 1144, then a stopping condition has been reached, and workflow orchestrator 710 can return the labeled result, including the confidence determined by QMS 750 for the label (step 1146))
Claim 7
Prendki in view of Martin discloses all the elements as stated in Claim 6 and further discloses wherein determining the target pre-filter procedure according to the task type of the target labeling task comprises: obtaining a set of pre-filter procedures, wherein each pre-filter procedure in the set of pre-filter procedures corresponds to a task type, and each pre-filter procedure includes discarding negatives in the manual labeling data; matching the task type of the target labeling task with one or more task types corresponding to the set of pre-filter procedures to identifying a matched pre-filter procedure in the set of pre-filter procedures; assigning the matched pre-filter procedure as the target pre-filter procedure (Prendki, ¶¶ [0065], [0067], and [0110] many labeling tasks require specialized knowledge in particular technologies or languages (e.g., computer vision, translation) or access to a specific kind of tools (e.g., software or hardware for Light Detection and Ranging, or LiDAR); some of these specialized or smaller labeling providers are able to provide superior labeling accuracy and faster tum-around time for certain types of tasks than their larger counterparts because they employ annotators that are more capable, knowledgeable, and skilled in handling those types of tasks; allowing customers to rely on a computer-implemented labeling marketplace to find the labeling provider that is most likely to produce the highest labeling quality and in a timely-manner, regardless of the size of the labeling providers and for whatever types of labeling tasks; determine the minimum criteria required for handling the labeling tasks requested by the customers; labeling partners unable to perform the labeling task because of an inability to perform tasks with the data type and task type requested may be eliminated from consideration for the labeling task; the minimum criteria may be determined during one of the initial phases of the system's processes so any labeling partners that are not capable of handling the requested tasks could be eliminated from the selection pool; ¶¶ [0191]-[0200] with FIG. 8: Hybrid Labeling Marketplace: find awesome labeling partners that are available immediately; process flow: 1. tell us about your project, expectations, and time and budgetary constraints; 2. choose between the recommended auto-labeling and manual labeling options; Mix-and-Match; 3. you get a precise estimate of the cost, time and accuracy; once the task launches, your data is immediately being labeled; 4. use the human-in-the-loop module to find mislabeled data and send it for revision either to the same partner or a different partner; human-in-the-loop module: visualize, audit, and fix your labels and measure labeling quality reliably; 1. upload your existing labels or select a labeling task after the labels are sent back by your labeling partner; 2. view label anomalies or query data using your own validation criteria; no need to ever review an entire dataset anymore; 3. fix faulty labels yourself with an annotation tool or send them back to either the same or a different labeling partner; 4. as you perfect your labels, they are automatically versioned, so that you can revert to an older version if needed; ¶ [0209]: if the system determines that a labeling task includes an amount of data that is too large for any one of the labeling partners in the selection pool, the system may assign various portions of the task to multiple labeling partners; the system may also assign a labeling task to several labeling partners if it determines that diversity among annotators may produce higher labeling quality) (Martin, ¶¶ [0055]-[0068] with FIGS. 2-3: Labelers (including labelers of different types) can be composed together into directed graphs as needed, such that each individual labeler solves a portion of an overall classification problem, and the results are aggregated together to form the overall labeled output; Labeling platform 104 may include multiple types of labelers and multiple labelers of each type; translation by conditioning layer 304 may be required because the data domain external to the kernel core logic 302 may be different than the kernel's data domain; the external data domain may be use-case specific and technology agnostic, while the kernel's data domain may be technology-specific and use case agnostic; the conditioning layer 304 may also perform validation on inbound data; e.g., for one use case, a solid black image may be valid for training/inferring, while for other use cases, it may not; if it is not, the conditioning layer 304 may, for example, include a filter to remove solid black images; alternatively, it might reject such input and issue an exception output; ¶¶ [0095]-[0122] with FIGS. 7 and 8A-D: ML labelers, human labelers and other labelers can be combined into a confidence-driven workflow (CDW); a CDW can thus be considered a labeler that encapsulates a collection of other labelers, and more particularly, a collection of labelers of the same arity; the encapsulated labelers can be consulted in sequence and their individual results incorporated into an overall result until a configured threshold confidence target for an overall result is reached; a labeling request may be received as a workflow task with accompanying task information, such as a task type description and constraints; examples of constraints include, but are not limited to, cost constraints in one or more dimensions (e.g., time limit, monetary limit), target threshold confidence or other constraints; CDW 700 encapsulates ML labeler 712, blind judgement human labeler 714, and open judgement human labeler 716, though it should be appreciated that a CDW can encapsulate any number of labelers of various types; each labeler instance may have a labeler instance score determined, for example, by QMS 750, and that corresponds to the probability that the labeler instance will produce an accurate label for a given task; the QMS 750 can score how often the labeler instance was correct when it labeled images as "tumor" and score how often the labeler instance was correct when it labeled images as "no tumor"; the answer-specific scores can be used in determining confidence estimates for an actual result output by the labeler instance; human labeler instance scores and ML labeler instance scores can be tied to specific labeling task types in the scoring system; for a labeler instance that can produce labels for multiple task types, QMS 750 may determine scores on the different task types as skill scores to differentiate the labeler instance's performance across different labeling task types; a labeler instance may have an associated response time (temporal cost) that is an estimate of how long it will take that labeler instance to perform a task; a labeler instance may have an associated price (monetary cost) that is an estimate of the price for that labeler instance to perform a task; for a particular labeling task, the workflow orchestrator 710 may select from labelers suited for that type of task; a labeler may also have a labeler score; a labeler's score corresponds to the probability that a labeler will produce an accurate label for a given task; a workflow orchestrator 710 to dynamically determine an execution path for processing a labeling request ( question) to produce a final labeled result based on confidence and cost constraint configuration; workflow orchestrator 710 uses the task information for the workflow task, the labeler descriptions, possibly dynamic labeler characteristics such as cost, availability, and timeliness, and quality metrics from QMS 750 to dynamically determine a path through the constituent labelers 712, 714, 716 to produce a labeled output for the task which satisfies a configured target confidence threshold (a minimum confidence threshold); the path may also be selected to minimize costs in one or more dimensions; workflow orchestrator 710 uses the labeler scores and costs associated with the labelers to determine one or more viable paths through the labelers; once a viable path is determined, workflow orchestrator 710 may route the labeling request through that path until a result reaches a threshold confidence target or the path is exhausted; the path may be re-evaluated and changed at any time based on the actual results produced by each consultation; the order of consultation of constituent labelers may proceed from least expensive labeler to most expensive labeler, for example, in an attempt to reach the confidence threshold target with the least cost; if ML labeler 712 is the least expensive labeler, blind judgement human labeler 714 is more expensive because labeling requests to blind judgement human labeler 714 are routed to human specialists who have a higher monetary cost based on compensation for work performed, and human labeler 716 is the most expensive of the constituent labelers because labeling requests to human labeler 716 are routed to human specialists with higher expertise and compensation levels than the human specialists associated with blind judgement human labeler 714, then workflow orchestrator 710 may favor consulting ML labeler 712 first, then human labeler 714, then human labeler 716; in the example of FIG. 8B, the labeling request is routed to ML labeler 712, which returns an answer for which a confidence estimate 806 is determined; the labeling request is then routed to blind judgement human labeler 714, which produces an answer for which a confidence estimate 808 is determined; as the target confidence threshold has not been reached, the labeling request can be routed to blind judgement human labeler 714 twice because there are labeler instances (human specialists 727) remaining in the pool of human labeler 714 that have not yet been consulted for the labeling request; the confidence estimate 810 for the answer produced by the second consultation with human labeler 714 may incorporate confidence estimates 806, 808, however, exceeds the target confidence threshold; thus, that answer can be used for the final labeled results of CDW 700; ¶¶ [0181]-[0215] with FIGS. 11A-B: route tasks to labelers based on any number of constraints that match the labelers' descriptions; specific sequencing of labelers can be specified to achieve predefined workflows; at step 1102, workflow orchestrator 710 applies criteria to filter out labelers from consideration; if there are cost constraints (temporal, monetary or other cost) and/or confidence constraints, workflow orchestrator 710 performs a path search to find the optimal set of labelers (step 1114); each path endpoint can represent a specific path of labelers (that is, a sequence of labelers consulted) and QMS 750 estimates the a priori endpoint confidence for the path (step 1120) before the path is executed; if the path meets the task constraints based on the estimated total path monetary cost, total path time, and end-point confidence, the path can be added to a set of viable paths (step 1120); otherwise the path can be discarded (step 1122); if at least one viable path is found, workflow orchestrator can select a path from the one or more viable paths (step 1128); if the selected viable path includes only one labeler (one node) (as determined at step 1129), the minimum confidence required for the labeler in order to maintain viability of the planned path may be determined (step 1130); at step 1132, workflow orchestrator 710 sends the task to that labeler with a confidence constraint, the confidence constraint including the required minimum confidence determined by QMS 750; at any point, if all remaining constraints and optimizations are satisfied by multiple labelers, then a random selection is made from these labelers; workflow orchestrator 710 selects the first labeler in a selected path (step 1134) and sends the task to that labeler with or without a confidence constraint (step 1136); at step 1138, workflow orchestrator 710 receives an output of the labeler; workflow orchestrator 710 determines the confidence estimate for the label returned by the labeler (step 1142); the confidence estimate returned for the label by QMS 750 meets the confidence threshold target, as determined at step 1144, then a stopping condition has been reached, and workflow orchestrator 710 can return the labeled result, including the confidence determined by QMS 750 for the label (step 1146)).
Claim 8
Prendki in view of Martin discloses all the elements as stated in Claim 2 and further discloses wherein determining, based on the task type of the target labeling task, whether the existing automatic labeling procedure corresponding to the target labeling task is present comprises: obtaining a set of historical automatic labeling procedures, wherein each historical automatic labeling procedure of the set of historical automatic labeling procedures corresponds to a task type; and matching the task type of the target labeling task with task types corresponding to the set of historical automatic labeling procedures to identify a matched historical automatic labeling procedure in the set of historical automatic labeling procedures (Prendki, ¶ [0050]: efficiently compare the capabilities, availability, performance, and costs of various labeling providers; a computer-implemented labeling marketplace, through which a suitable labeling partner (i.e., labeling provider partnered with the computer-implemented labeling marketplace) will be automatically selected, based both on historical and real-time metrics, to handle the task requested by a customer; provide a short-list of labeling partners that are similarly determined to be suitable for handling the task; such selection processes may involve eliminating from the selection pool labeling partners that are not capable of handling the requested task; determining real-time availability of the labeling partners, and evaluating the labeling partners' performance history, costs, and various other factors that may be relevant to delivering on the expectations of the customers; incorporate a quality control process, in which the tasks completed by the labeling partners may be evaluated; ¶¶ [0056]-[0057], [0059], [0065], [0067],and [0110] with FIG. 3: recommendation computer system 306 is programmed using the real-time evaluation instructions 308 to query the provider metrics database 310 and to select, automatically and based upon programmed recommendation or evaluation algorithms, one or more of the label service providers 314, 316, 318 that is optimal to perform one or more labeling tasks on the user dataset; provides labeling services to customers by finding a labeling partner that is suitable to handle the tasks requested by the customers; a customer may rely on the system to automatically select the labeling partner that is most likely to provide the highest labeling quality in a timely manner, or alternatively, the system may provide the customer with a short-list of labeling partners determined to be suitable for handling the requested task; many labeling tasks require specialized knowledge in particular technologies or languages (e.g., computer vision, translation) or access to a specific kind of tools (e.g., software or hardware for Light Detection and Ranging, or LiDAR); some of these specialized or smaller labeling providers are able to provide superior labeling accuracy and faster tum-around time for certain types of tasks than their larger counterparts because they employ annotators that are more capable, knowledgeable, and skilled in handling those types of tasks; allowing customers to rely on a computer-implemented labeling marketplace to find the labeling provider that is most likely to produce the highest labeling quality and in a timely-manner, regardless of the size of the labeling providers and for whatever types of labeling tasks; every labeling task created in association to the customer's project may then automatically be recognized with the matching data type and task type; labeling partners unable to perform the labeling task because of an inability to perform tasks with the data type and task type requested may be eliminated from consideration for the labeling task; the minimum criteria may be determined during one of the initial phases of the system's processes so any labeling partners that are not capable of handling the requested tasks could be eliminated from the selection pool; ¶¶ [0146]-[0154]: utilize a multi-objective optimization process to find a labeling partner that is most likely to produce the highest quality of work; examples of the factors considered during the optimization process include, but are not limited to, the following: a) turn-around time (e.g., on average how long does the labeling partner need to complete and return a task); b) labeling quality (e.g., how accurate is the labeled data); c) average cost (e.g., the average cost charged by the labeling partner for various types of tasks); d) revenue (e.g., how much did the labeling partner earn for a given period of time); e) number of, and type of, tasks previously offered to the labeling partner; f) number of labeling partners within the computer-implemented labeling marketplace; and g) customer's preferences/criteria; the optimization process may be programmed using any of: the Dantzig simplex algorithm, extensions or variants thereof; combinatorial algorithms; or quantum optimization algorithms; ¶ [0168]: identify a labeling partner that is most suitable for handling the customer's task by considering, among others, real-time availability of the annotators employed by the various labeling partners; allows the system to track each of their annotator's computer activity, or to provide some other indicators to verify whether the annotators are available to receive a labeling task; ¶¶ [0169]-[0173] with FIG. 5F: to provide customers with consistent and efficient turn-around time for the delivery of the requested tasks, the system may place certain time constraints on the labeling partners; the system may provide the labeling partner a limited amount of time to complete any task accepted by the labeling partner, which time period may be determined based on the historical data of the labeling partner or other labeling partners (e.g., average amount of time needed to complete a task of a similar type); the customer may be shown a ranked list of labeling partners and be requested to select its top three choices; the first selected labeling partner may be offered the labeling task; if the first selected labeling partner fails to respond in time, the system may automatically route the request to the labeling partner that was the customer's second top choice; if all three of the customer's initial selections decline the request or fail to timely respond, the system may notify the customer of this failed status and may prompt the customer to continue to select additional labeling partners or to make adjustments to its labeling methodology until a suitable labeling partner accepts the offer for the labeling task; FIG. 5F illustrates a computer display device that is displaying an example of a graphical user interface with which a customer can select one or more labeling partners and view recommendations of labeling partners; ¶¶ [0176]-[0183] with FIG. 6B: select a labeling partner that is suitable for the customer's requested task in a sequential manner, such that the task is offered to only one labeling partner at a time; the system may re-evaluate all of the labeling partners within the selection pool, e.g., in situations where enough time has passed, and real-time metrics associated with the labeling partners have substantially changed; select a labeling partner that is suitable for the customer's requested task by employing a self-serve process, bidding process, or reverse-bidding process; if, e.g., a customer prioritizes cost as an important criteria for the task, the task may be offered to several labeling partners at the same time; in such a case, the labeling partner that provides the lowest cost for the task may be selected; alternatively, the system may provide to the customer a list of labeling providers determined to be suitable for the task, along with the costs associated with each of the providers; the system (which has historical data for price, time and accuracy from previous jobs, on a use case-by-use case basis), is programmed to compute the z-score of each valid labeling partner for price, time and accuracy; each of the z-scores is weighted with the number of points allocated to the feature (for price and time, a -1 factor is added, since lower price or time is better); the system is programmed to compute the sum and to return a ranked list of recommendations based on the sum values; FIG. 6B illustrates an example of a graphical user interface that can be programmed as presentation output of an ordered set of recommended labeling partners; provide task-level feedback that provides information about why a specific task was not offered to a labeling partner (e.g., insufficient availability of annotators, lack of tooling, etc.) or aggregated feedback that provides information about why specific types of tasks, or tasks in general, are not offered to the labeling partner (e.g., statistical insights about subpar turn-around time or labeling accuracy); ¶¶ [0208]-[0209]: select a labeling partner to handle the labeling at an experiment-level (i.e., the task in its entirety), which may involve one or more loops of labeling; this will provide consistency in the labeled data and minimize the inefficiencies that come from transitioning the task between several labeling partners; select several labeling partners to handle the task, each labeling partner being assigned to a particular loop of the task; e.g., if the system determines that a task can be partitioned into several loops and some of the loops are better suitable for one labeling partner while other loops are better suitable for another labeling partner, the system may assign the task to both labeling partners; if the system determines that a labeling task includes an amount of data that is too large for any one of the labeling partners in the selection pool, the system may assign various portions of the task to multiple labeling partners; the system may also assign a labeling task to several labeling partners if it determines that diversity among annotators may produce higher labeling quality; ¶¶ [0218]-[0223] with FIG. 1: to provide labeling services to a customer, the labeling partner needs to fulfill the following expectations: (a) possess a task force capable/trained to manage the particular type of task; (b) possess the proper annotation tool for the said task; be compliant to the level required by the user; (c) be compliant to the level required by the user; and (d) have at least one annotator available immediately; at step 102, the system applies the minimum criteria to each of the pre-selected labeling partners within the selection pool and eliminates any labeling partners that do not meet the minimum criteria, i.e., that are not capable of handling the task; at steps 103 and 104, the system applies an optimization process that evaluates real-time and historical metrics of all the labeling partners remaining in the selection pool; at step 103, check the real-time availability of the labeling partners, which may include checking real-time availability of the annotators employed by each of the labeling partners; at step 104, evaluates the historical metrics along with the real-time metrics; examples of the historical metrics evaluated by the system may include each of the labeling partners' average labeling accuracy, average turn-around time, and average pricing, both at an overall level and task-specific level (e.g., by grouping together similar types of tasks); based on the real-time and historical metrics, identify a labeling partner that is the most likely to produce the highest quality of work and in a timely-manner; at step 109, if a labeling partner either refuses the task or does not reply within a set amount of time, remove the labeling partner from the selection pool and identifies another labeling partner that is suitable for handling the task; identify the subsequent labeling partner by re-evaluating the real-time and historical metrics of the labeling partners left in the selection pool; alternatively, identify the subsequent labeling partner based on the previous evaluation, e.g., in situations where insignificant amount of time has passed since the previous evaluation; at step 113, if the system determines that the quality of the labeling partner's work is acceptable, the labeling task is deemed complete and the labeled data is provided back to the customer (e.g., to a data storage dedicated to the customer); the system may store the labeled data in a historical database and index the data with a unique identification assigned to the labeling partner; the system may also index the data based on the type of data labeled, the type of the labeling task, and the quality of the labeled data (e.g., labeling accuracy, turn-around time.); the historical metrics evaluated in step 104 may correspond to the information stored in the historical database; ¶ [0229] with FIGS. 1-2: the steps of identifying/selecting an optimal labeling partner (i.e., steps 102-108 within brackets 260) may be skipped for loops n> 1, if the system determines that the same labeling partner should be used throughout the experiment; alternatively, the steps within brackets 260 may be kept in place if the system determines that different labeling partners should be used within the same experiment) (Martin, ¶¶ [0085]-[0094] with FIG. 6: the ML labeler may be implemented as a wrapper for an ML model on an ML platform 650 running locally or on a remote ML platform system; the ML labeler configuration can specify an ML algorithm to use and, based on the ML algorithm specified, labeling platform 104 configures the labeler with the code to connect to the appropriate ML platform 650 to train and use the specified ML algorithm; ML labeler 600 includes training component 615 executable to train an ML algorithm; the training component 615 includes an experiment coordinator 616 that interfaces with ML platform 650 to train multiple challenger models ( e.g., using various hyperparameters or other mechanisms for training multiple candidate models known or developed in the art) and a challenger model evaluator 618 that evaluates candidate ML models against each other and the current active model to determine which should be the current active model for inferring answers to labeling requests; the output is a champion ML model that represents the best model currently producible given the available training data; the training component 615 thus determines the ML model to use as the current active model for inferring answers to labeling requests; training triggers may be based on, for example, an amount of training data received by the labeler, quality metrics received by the labeler, elapsed time, or other criteria).
Claim 9
Prendki in view of Martin discloses all the elements as stated in Claim 8 and further discloses wherein determining, based on the task type of the target labeling task, whether an existing automatic labeling procedure corresponding to the target labeling task is present further comprises: if no matched historical automatic labeling procedure is identified, assigning the automatic labeling data to the manual labeling data (Prendki, ¶¶ [0065], [0067], and [0110] many labeling tasks require specialized knowledge in particular technologies or languages (e.g., computer vision, translation) or access to a specific kind of tools (e.g., software or hardware for Light Detection and Ranging, or LiDAR); some of these specialized or smaller labeling providers are able to provide superior labeling accuracy and faster tum-around time for certain types of tasks than their larger counterparts because they employ annotators that are more capable, knowledgeable, and skilled in handling those types of tasks; allowing customers to rely on a computer-implemented labeling marketplace to find the labeling provider that is most likely to produce the highest labeling quality and in a timely-manner, regardless of the size of the labeling providers and for whatever types of labeling tasks; determine the minimum criteria required for handling the labeling tasks requested by the customers; labeling partners unable to perform the labeling task because of an inability to perform tasks with the data type and task type requested may be eliminated from consideration for the labeling task; the minimum criteria may be determined during one of the initial phases of the system's processes so any labeling partners that are not capable of handling the requested tasks could be eliminated from the selection pool; ¶¶ [0191]-[0200] with FIG. 8: Hybrid Labeling Marketplace: find awesome labeling partners that are available immediately; process flow: 1. tell us about your project, expectations, and time and budgetary constraints; 2. choose between the recommended auto-labeling and manual labeling options; Mix-and-Match; 3. you get a precise estimate of the cost, time and accuracy; once the task launches, your data is immediately being labeled; 4. use the human-in-the-loop module to find mislabeled data and send it for revision either to the same partner or a different partner; human-in-the-loop module: visualize, audit, and fix your labels and measure labeling quality reliably; 1. upload your existing labels or select a labeling task after the labels are sent back by your labeling partner; 2. view label anomalies or query data using your own validation criteria; no need to ever review an entire dataset anymore; 3. fix faulty labels yourself with an annotation tool or send them back to either the same or a different labeling partner; 4. as you perfect your labels, they are automatically versioned, so that you can revert to an older version if needed; ¶ [0209]: if the system determines that a labeling task includes an amount of data that is too large for any one of the labeling partners in the selection pool, the system may assign various portions of the task to multiple labeling partners; the system may also assign a labeling task to several labeling partners if it determines that diversity among annotators may produce higher labeling quality) (Martin, ¶¶ [0055]-[0068] with FIGS. 2-3: Labelers (including labelers of different types) can be composed together into directed graphs as needed, such that each individual labeler solves a portion of an overall classification problem, and the results are aggregated together to form the overall labeled output; Labeling platform 104 may include multiple types of labelers and multiple labelers of each type; translation by conditioning layer 304 may be required because the data domain external to the kernel core logic 302 may be different than the kernel's data domain; the external data domain may be use-case specific and technology agnostic, while the kernel's data domain may be technology-specific and use case agnostic; the conditioning layer 304 may also perform validation on inbound data; e.g., for one use case, a solid black image may be valid for training/inferring, while for other use cases, it may not; if it is not, the conditioning layer 304 may, for example, include a filter to remove solid black images; alternatively, it might reject such input and issue an exception output; ¶¶ [0095]-[0122] with FIGS. 7 and 8A-D: ML labelers, human labelers and other labelers can be combined into a confidence-driven workflow (CDW); a CDW can thus be considered a labeler that encapsulates a collection of other labelers, and more particularly, a collection of labelers of the same arity; the encapsulated labelers can be consulted in sequence and their individual results incorporated into an overall result until a configured threshold confidence target for an overall result is reached; a labeling request may be received as a workflow task with accompanying task information, such as a task type description and constraints; examples of constraints include, but are not limited to, cost constraints in one or more dimensions (e.g., time limit, monetary limit), target threshold confidence or other constraints; CDW 700 encapsulates ML labeler 712, blind judgement human labeler 714, and open judgement human labeler 716, though it should be appreciated that a CDW can encapsulate any number of labelers of various types; each labeler instance may have a labeler instance score determined, for example, by QMS 750, and that corresponds to the probability that the labeler instance will produce an accurate label for a given task; the QMS 750 can score how often the labeler instance was correct when it labeled images as "tumor" and score how often the labeler instance was correct when it labeled images as "no tumor"; the answer-specific scores can be used in determining confidence estimates for an actual result output by the labeler instance; human labeler instance scores and ML labeler instance scores can be tied to specific labeling task types in the scoring system; for a labeler instance that can produce labels for multiple task types, QMS 750 may determine scores on the different task types as skill scores to differentiate the labeler instance's performance across different labeling task types; a labeler instance may have an associated response time (temporal cost) that is an estimate of how long it will take that labeler instance to perform a task; a labeler instance may have an associated price (monetary cost) that is an estimate of the price for that labeler instance to perform a task; for a particular labeling task, the workflow orchestrator 710 may select from labelers suited for that type of task; a labeler may also have a labeler score; a labeler's score corresponds to the probability that a labeler will produce an accurate label for a given task; a workflow orchestrator 710 to dynamically determine an execution path for processing a labeling request ( question) to produce a final labeled result based on confidence and cost constraint configuration; workflow orchestrator 710 uses the task information for the workflow task, the labeler descriptions, possibly dynamic labeler characteristics such as cost, availability, and timeliness, and quality metrics from QMS 750 to dynamically determine a path through the constituent labelers 712, 714, 716 to produce a labeled output for the task which satisfies a configured target confidence threshold (a minimum confidence threshold); the path may also be selected to minimize costs in one or more dimensions; workflow orchestrator 710 uses the labeler scores and costs associated with the labelers to determine one or more viable paths through the labelers; once a viable path is determined, workflow orchestrator 710 may route the labeling request through that path until a result reaches a threshold confidence target or the path is exhausted; the path may be re-evaluated and changed at any time based on the actual results produced by each consultation; the order of consultation of constituent labelers may proceed from least expensive labeler to most expensive labeler, for example, in an attempt to reach the confidence threshold target with the least cost; if ML labeler 712 is the least expensive labeler, blind judgement human labeler 714 is more expensive because labeling requests to blind judgement human labeler 714 are routed to human specialists who have a higher monetary cost based on compensation for work performed, and human labeler 716 is the most expensive of the constituent labelers because labeling requests to human labeler 716 are routed to human specialists with higher expertise and compensation levels than the human specialists associated with blind judgement human labeler 714, then workflow orchestrator 710 may favor consulting ML labeler 712 first, then human labeler 714, then human labeler 716; in the example of FIG. 8B, the labeling request is routed to ML labeler 712, which returns an answer for which a confidence estimate 806 is determined; the labeling request is then routed to blind judgement human labeler 714, which produces an answer for which a confidence estimate 808 is determined; as the target confidence threshold has not been reached, the labeling request can be routed to blind judgement human labeler 714 twice because there are labeler instances (human specialists 727) remaining in the pool of human labeler 714 that have not yet been consulted for the labeling request; the confidence estimate 810 for the answer produced by the second consultation with human labeler 714 may incorporate confidence estimates 806, 808, however, exceeds the target confidence threshold; thus, that answer can be used for the final labeled results of CDW 700; ¶¶ [0181]-[0215] with FIGS. 11A-B: route tasks to labelers based on any number of constraints that match the labelers' descriptions; specific sequencing of labelers can be specified to achieve predefined workflows; at step 1102, workflow orchestrator 710 applies criteria to filter out labelers from consideration; if there are cost constraints (temporal, monetary or other cost) and/or confidence constraints, workflow orchestrator 710 performs a path search to find the optimal set of labelers (step 1114); each path endpoint can represent a specific path of labelers (that is, a sequence of labelers consulted) and QMS 750 estimates the a priori endpoint confidence for the path (step 1120) before the path is executed; if the path meets the task constraints based on the estimated total path monetary cost, total path time, and end-point confidence, the path can be added to a set of viable paths (step 1120); otherwise the path can be discarded (step 1122); if at least one viable path is found, workflow orchestrator can select a path from the one or more viable paths (step 1128); if the selected viable path includes only one labeler (one node) (as determined at step 1129), the minimum confidence required for the labeler in order to maintain viability of the planned path may be determined (step 1130); at step 1132, workflow orchestrator 710 sends the task to that labeler with a confidence constraint, the confidence constraint including the required minimum confidence determined by QMS 750; at any point, if all remaining constraints and optimizations are satisfied by multiple labelers, then a random selection is made from these labelers; workflow orchestrator 710 selects the first labeler in a selected path (step 1134) and sends the task to that labeler with or without a confidence constraint (step 1136); at step 1138, workflow orchestrator 710 receives an output of the labeler; workflow orchestrator 710 determines the confidence estimate for the label returned by the labeler (step 1142); the confidence estimate returned for the label by QMS 750 meets the confidence threshold target, as determined at step 1144, then a stopping condition has been reached, and workflow orchestrator 710 can return the labeled result, including the confidence determined by QMS 750 for the label (step 1146)).
Claim 12
Prendki in view of Martin discloses all the elements as stated in Claim 2 and further discloses wherein determining whether the existing automatic labeling procedure satisfies the first quality evaluation index comprises: if the existing automatic labeling procedure satisfies the first quality evaluation index, labeling the automatic labeling data with the existing automatic labeling procedure (Prendki, ¶ [0050]: efficiently compare the capabilities, availability, performance, and costs of various labeling providers; a computer-implemented labeling marketplace, through which a suitable labeling partner (i.e., labeling provider partnered with the computer-implemented labeling marketplace) will be automatically selected, based both on historical and real-time metrics, to handle the task requested by a customer; provide a short-list of labeling partners that are similarly determined to be suitable for handling the task; such selection processes may involve eliminating from the selection pool labeling partners that are not capable of handling the requested task; determining real-time availability of the labeling partners, and evaluating the labeling partners' performance history, costs, and various other factors that may be relevant to delivering on the expectations of the customers; incorporate a quality control process, in which the tasks completed by the labeling partners may be evaluated; ¶¶ [0056]-[0057], [0059], [0065], [0067],and [0110] with FIG. 3: recommendation computer system 306 is programmed using the real-time evaluation instructions 308 to query the provider metrics database 310 and to select, automatically and based upon programmed recommendation or evaluation algorithms, one or more of the label service providers 314, 316, 318 that is optimal to perform one or more labeling tasks on the user dataset; provides labeling services to customers by finding a labeling partner that is suitable to handle the tasks requested by the customers; a customer may rely on the system to automatically select the labeling partner that is most likely to provide the highest labeling quality in a timely manner, or alternatively, the system may provide the customer with a short-list of labeling partners determined to be suitable for handling the requested task; many labeling tasks require specialized knowledge in particular technologies or languages (e.g., computer vision, translation) or access to a specific kind of tools (e.g., software or hardware for Light Detection and Ranging, or LiDAR); some of these specialized or smaller labeling providers are able to provide superior labeling accuracy and faster tum-around time for certain types of tasks than their larger counterparts because they employ annotators that are more capable, knowledgeable, and skilled in handling those types of tasks; allowing customers to rely on a computer-implemented labeling marketplace to find the labeling provider that is most likely to produce the highest labeling quality and in a timely-manner, regardless of the size of the labeling providers and for whatever types of labeling tasks; labeling partners unable to perform the labeling task because of an inability to perform tasks with the data type and task type requested may be eliminated from consideration for the labeling task; the minimum criteria may be determined during one of the initial phases of the system's processes so any labeling partners that are not capable of handling the requested tasks could be eliminated from the selection pool; ¶¶ [0146]-[0154]: utilize a multi-objective optimization process to find a labeling partner that is most likely to produce the highest quality of work; examples of the factors considered during the optimization process include, but are not limited to, the following: a) turn-around time (e.g., on average how long does the labeling partner need to complete and return a task); b) labeling quality (e.g., how accurate is the labeled data); c) average cost (e.g., the average cost charged by the labeling partner for various types of tasks); d) revenue (e.g., how much did the labeling partner earn for a given period of time); e) number of, and type of, tasks previously offered to the labeling partner; f) number of labeling partners within the computer-implemented labeling marketplace; and g) customer's preferences/criteria; the optimization process may be programmed using any of: the Dantzig simplex algorithm, extensions or variants thereof; combinatorial algorithms; or quantum optimization algorithms; ¶ [0168]: identify a labeling partner that is most suitable for handling the customer's task by considering, among others, real-time availability of the annotators employed by the various labeling partners; allows the system to track each of their annotator's computer activity, or to provide some other indicators to verify whether the annotators are available to receive a labeling task; ¶¶ [0169]-[0173] with FIG. 5F: to provide customers with consistent and efficient turn-around time for the delivery of the requested tasks, the system may place certain time constraints on the labeling partners; the system may provide the labeling partner a limited amount of time to complete any task accepted by the labeling partner, which time period may be determined based on the historical data of the labeling partner or other labeling partners (e.g., average amount of time needed to complete a task of a similar type); the customer may be shown a ranked list of labeling partners and be requested to select its top three choices; the first selected labeling partner may be offered the labeling task; if the first selected labeling partner fails to respond in time, the system may automatically route the request to the labeling partner that was the customer's second top choice; if all three of the customer's initial selections decline the request or fail to timely respond, the system may notify the customer of this failed status and may prompt the customer to continue to select additional labeling partners or to make adjustments to its labeling methodology until a suitable labeling partner accepts the offer for the labeling task; FIG. 5F illustrates a computer display device that is displaying an example of a graphical user interface with which a customer can select one or more labeling partners and view recommendations of labeling partners; ¶¶ [0176]-[0183] with FIG. 6B: select a labeling partner that is suitable for the customer's requested task in a sequential manner, such that the task is offered to only one labeling partner at a time; the system may re-evaluate all of the labeling partners within the selection pool, e.g., in situations where enough time has passed, and real-time metrics associated with the labeling partners have substantially changed; select a labeling partner that is suitable for the customer's requested task by employing a self-serve process, bidding process, or reverse-bidding process; if, e.g., a customer prioritizes cost as an important criteria for the task, the task may be offered to several labeling partners at the same time; in such a case, the labeling partner that provides the lowest cost for the task may be selected; alternatively, the system may provide to the customer a list of labeling providers determined to be suitable for the task, along with the costs associated with each of the providers; the system (which has historical data for price, time and accuracy from previous jobs, on a use case-by-use case basis), is programmed to compute the z-score of each valid labeling partner for price, time and accuracy; each of the z-scores is weighted with the number of points allocated to the feature (for price and time, a -1 factor is added, since lower price or time is better); the system is programmed to compute the sum and to return a ranked list of recommendations based on the sum values; FIG. 6B illustrates an example of a graphical user interface that can be programmed as presentation output of an ordered set of recommended labeling partners; provide task-level feedback that provides information about why a specific task was not offered to a labeling partner (e.g., insufficient availability of annotators, lack of tooling, etc.) or aggregated feedback that provides information about why specific types of tasks, or tasks in general, are not offered to the labeling partner (e.g., statistical insights about subpar turn-around time or labeling accuracy); ¶¶ [0208]-[0209]: select a labeling partner to handle the labeling at an experiment-level (i.e., the task in its entirety), which may involve one or more loops of labeling; this will provide consistency in the labeled data and minimize the inefficiencies that come from transitioning the task between several labeling partners; select several labeling partners to handle the task, each labeling partner being assigned to a particular loop of the task; e.g., if the system determines that a task can be partitioned into several loops and some of the loops are better suitable for one labeling partner while other loops are better suitable for another labeling partner, the system may assign the task to both labeling partners; if the system determines that a labeling task includes an amount of data that is too large for any one of the labeling partners in the selection pool, the system may assign various portions of the task to multiple labeling partners; the system may also assign a labeling task to several labeling partners if it determines that diversity among annotators may produce higher labeling quality; ¶¶ [0218]-[0221] with FIG. 1: to provide labeling services to a customer, the labeling partner needs to fulfill the following expectations: (a) possess a task force capable/trained to manage the particular type of task; (b) possess the proper annotation tool for the said task; be compliant to the level required by the user; (c) be compliant to the level required by the user; and (d) have at least one annotator available immediately; at step 102, the system applies the minimum criteria to each of the pre-selected labeling partners within the selection pool and eliminates any labeling partners that do not meet the minimum criteria, i.e., that are not capable of handling the task; at steps 103 and 104, the system applies an optimization process that evaluates real-time and historical metrics of all the labeling partners remaining in the selection pool; at step 103, check the real-time availability of the labeling partners, which may include checking real-time availability of the annotators employed by each of the labeling partners; at step 104, evaluates the historical metrics along with the real-time metrics; examples of the historical metrics evaluated by the system may include each of the labeling partners' average labeling accuracy, average turn-around time, and average pricing, both at an overall level and task-specific level (e.g., by grouping together similar types of tasks); based on the real-time and historical metrics, identify a labeling partner that is the most likely to produce the highest quality of work and in a timely-manner; at step 109, if a labeling partner either refuses the task or does not reply within a set amount of time, remove the labeling partner from the selection pool and identifies another labeling partner that is suitable for handling the task; identify the subsequent labeling partner by re-evaluating the real-time and historical metrics of the labeling partners left in the selection pool; alternatively, identify the subsequent labeling partner based on the previous evaluation, e.g., in situations where insignificant amount of time has passed since the previous evaluation; ¶ [0229] with FIGS. 1-2: the steps of identifying/selecting an optimal labeling partner (i.e., steps 102-108 within brackets 260) may be skipped for loops n> 1, if the system determines that the same labeling partner should be used throughout the experiment; alternatively, the steps within brackets 260 may be kept in place if the system determines that different labeling partners should be used within the same experiment) (Martin, ¶¶ [0095]-[0122] with FIGS. 7 and 8A-D: ML labelers, human labelers and other labelers can be combined into a confidence-driven workflow (CDW); a CDW can thus be considered a labeler that encapsulates a collection of other labelers, and more particularly, a collection of labelers of the same arity; the encapsulated labelers can be consulted in sequence and their individual results incorporated into an overall result until a configured threshold confidence target for an overall result is reached; a labeling request may be received as a workflow task with accompanying task information, such as a task type description and constraints; examples of constraints include, but are not limited to, cost constraints in one or more dimensions (e.g., time limit, monetary limit), target threshold confidence or other constraints; CDW 700 encapsulates ML labeler 712, blind judgement human labeler 714, and open judgement human labeler 716, though it should be appreciated that a CDW can encapsulate any number of labelers of various types; each labeler instance may have a labeler instance score determined, for example, by QMS 750, and that corresponds to the probability that the labeler instance will produce an accurate label for a given task; the QMS 750 can score how often the labeler instance was correct when it labeled images as "tumor" and score how often the labeler instance was correct when it labeled images as "no tumor"; the answer-specific scores can be used in determining confidence estimates for an actual result output by the labeler instance; human labeler instance scores and ML labeler instance scores can be tied to specific labeling task types in the scoring system; for a labeler instance that can produce labels for multiple task types, QMS 750 may determine scores on the different task types as skill scores to differentiate the labeler instance's performance across different labeling task types; a labeler instance may have an associated response time (temporal cost) that is an estimate of how long it will take that labeler instance to perform a task; a labeler instance may have an associated price (monetary cost) that is an estimate of the price for that labeler instance to perform a task; for a particular labeling task, the workflow orchestrator 710 may select from labelers suited for that type of task; a labeler may also have a labeler score; a labeler's score corresponds to the probability that a labeler will produce an accurate label for a given task; a workflow orchestrator 710 to dynamically determine an execution path for processing a labeling request ( question) to produce a final labeled result based on confidence and cost constraint configuration; workflow orchestrator 710 uses the task information for the workflow task, the labeler descriptions, possibly dynamic labeler characteristics such as cost, availability, and timeliness, and quality metrics from QMS 750 to dynamically determine a path through the constituent labelers 712, 714, 716 to produce a labeled output for the task which satisfies a configured target confidence threshold (a minimum confidence threshold); the path may also be selected to minimize costs in one or more dimensions; workflow orchestrator 710 uses the labeler scores and costs associated with the labelers to determine one or more viable paths through the labelers; once a viable path is determined, workflow orchestrator 710 may route the labeling request through that path until a result reaches a threshold confidence target or the path is exhausted; the path may be re-evaluated and changed at any time based on the actual results produced by each consultation; the order of consultation of constituent labelers may proceed from least expensive labeler to most expensive labeler, for example, in an attempt to reach the confidence threshold target with the least cost; if ML labeler 712 is the least expensive labeler, blind judgement human labeler 714 is more expensive because labeling requests to blind judgement human labeler 714 are routed to human specialists who have a higher monetary cost based on compensation for work performed, and human labeler 716 is the most expensive of the constituent labelers because labeling requests to human labeler 716 are routed to human specialists with higher expertise and compensation levels than the human specialists associated with blind judgement human labeler 714, then workflow orchestrator 710 may favor consulting ML labeler 712 first, then human labeler 714, then human labeler 716; in the example of FIG. 8B, the labeling request is routed to ML labeler 712, which returns an answer for which a confidence estimate 806 is determined; the labeling request is then routed to blind judgement human labeler 714, which produces an answer for which a confidence estimate 808 is determined; as the target confidence threshold has not been reached, the labeling request can be routed to blind judgement human labeler 714 twice because there are labeler instances (human specialists 727) remaining in the pool of human labeler 714 that have not yet been consulted for the labeling request; the confidence estimate 810 for the answer produced by the second consultation with human labeler 714 may incorporate confidence estimates 806, 808, however, exceeds the target confidence threshold; thus, that answer can be used for the final labeled results of CDW 700).
Claim 14
Prendki in view of Martin discloses all the elements as stated in Claim 3 and further discloses wherein computing labeling quality and labeling metric of each automatic labeling procedure in the set of automatic labeling procedures comprises: extracting quality inspection data from the labeling data, the quality inspection data is a part of the labeling data; labeling the quality inspection data respectively with each automatic labeling procedure in the set of automatic labeling procedures; computing labeling quality of each automatic labeling procedure in the set of the automatic labeling procedures based on a result of labeling the quality inspection data by each automatic labeling procedure in the set of the automatic labeling procedures and the result of manual labeling the quality inspection data; obtaining task characteristics of the target labeling task, the task characteristics including quantity of data items of the automatic labeling data and an average cost for manually labeling a single data item of the automatic labeling data; computing labeling metric of each automatic labeling procedure in the set of the automatic labeling procedures respectively based on the task characteristics and the result of labeling the quality inspection data by each of the automatic labeling procedures (Prendki, ¶¶ [0114]-[0154]: collect various information of the labeling partners, which may be used by the optimization system to find the most suitable labeling partner for a customer-requested task; examples of such information that the system collects of the labeling partners may include the following: list of the types of tasks that the partner can deliver on (depends on tooling and labor workforce); available clearance; annotator information (e.g., type of tasks annotators are capable of working on); historical performance of the annotators (e.g., speed and accuracy); desired workload of annotators; and real-time availability of annotators; compute and maintain various metrics of the labeling providers; examples of such metrics may include a) average labeling quality (e.g., accuracy) for various types of tasks; b) average labeling speed for various types of tasks; c) typical availability; d) average price, and deviation from standard price; e) time since last offered task; f) time of idle (e.g., no assigned task); g) time taken to resolve labeling issues identified by the quality control process; and number of iterations required to solve a labeling issue; h) response time to the system's inquiry/notifications; and i) acceptance/rejection/ignore statistics of labeling tasks; utilize a multi-objective optimization process to find a labeling partner that is most likely to produce the highest quality of work; examples of the factors considered during the optimization process include, but are not limited to, the following: a) turn-around time (e.g., on average how long does the labeling partner need to complete and return a task); b) labeling quality (e.g., how accurate is the labeled data); c) average cost (e.g., the average cost charged by the labeling partner for various types of tasks); d) revenue (e.g., how much did the labeling partner earn for a given period of time); e) number of, and type of, tasks previously offered to the labeling partner; f) number of labeling partners within the computer-implemented labeling marketplace; and g) customer's preferences/criteria; the optimization process may be programmed using any of: the Dantzig simplex algorithm, extensions or variants thereof; combinatorial algorithms; or quantum optimization algorithms; Claim 10: the identifying by multi-objective optimization based on: an average time required by the labeling partners to complete and return a task; an average labeling accuracy of the labeling partners; an average price of the labeling partners; an average revenue of the labeling partners over a given period of time; a number of tasks previously offered to the labeling partners; and types of tasks previously offered to the labeling partners; ¶¶ [0186]-[0189]: verify the quality of the labeling tasks completed by the labeling partners by evaluating the completed tasks; the recommendation computer system 306 can be programmed with an anomaly detection engine that creates a distribution of the annotations as returned by the labeling partner and looks for potential anomalous annotations such as out of range values; the anomaly detection engine can be programmed for generating a z-score distribution in various features of interest; recommendation computer system 306 can be programmed with an auto-labeling/pre-labeling process, using a pre-trained machine learning model to generate synthetic generation, which are checked against the manual annotations provided by the labeling partner; in ML-driven data curation, which is active learning-based; data is selected dynamically during a training process; the training process ends when the user reaches their budget or the remaining data contains no new relevant information; ¶¶ [0060], [0174], and [0226]-[0229] with FIGS. 1-2: at step 201, a customer uploads to the system labeling instructions, data to be labeled, and a model to be trained using the data; at step 202, the customer decides a labeling approach (e.g., ML-based approach, manual approach, or self-labeling approach); at step 203, the curation process begins, and at step 204, the customer's model and data are sent to the data curation engine; the system's data preparation and data curation services employ an Active Learning approach where portions of the data are prioritized so the portions that are most effective in training the model are labeled first; the curation engine analyzes the data and identifies a first batch of data that it believes will be effective in training the data; the first batch of data is provided to a labeling partner in a similar fashion as the processes described in FIG. 1; once the labeling partner completes labeling the first batch of data, the labeled data is used to train the customer's model, as described in step 206; then, the model's performance is evaluated; at steps 207 and 208, if the system determines that the model's performance is acceptable, the system deems the model as sufficiently trained and provides the trained model and labeled data back to the customer (i.e., the experiment ends); the experiment may also end if the system determines that the customer's budget has ran out or if the model has reached its optimal performance; at step 205, if the system determines that the model's performance is unacceptable, the next loop of the experiment begins; going back to step 204, the data curation engine identifies the next batch of data and provides this batch to the labeling partner along with all of the previous batches of data; once the labeling partner completes labeling a batch of the data, the experiment continues until the model's performance reaches an acceptable level, the budget is reached, or optimal performance has been achieved because the remaining data is determined as redundant or irrelevant for training the model; ¶¶ [0203]-[0204]: provide quality assurances by identifying problematic labels; employ an adversarial process that involves providing deceptive or faulty data and evaluating the response of a service provider to receiving the deceptive or faulty data; provide the data labeled by one labeling partner to another labeling partner for cross-validation) (Martin, ¶¶ [0095]-[0147] with FIGS. 7 and 8A-D: ML labelers, human labelers and other labelers can be combined into a confidence-driven workflow (CDW); a CDW can thus be considered a labeler that encapsulates a collection of other labelers, and more particularly, a collection of labelers of the same arity; the encapsulated labelers can be consulted in sequence and their individual results incorporated into an overall result until a configured threshold confidence target for an overall result is reached; a labeling request may be received as a workflow task with accompanying task information, such as a task type description and constraints; examples of constraints include, but are not limited to, cost constraints in one or more dimensions (e.g., time limit, monetary limit), target threshold confidence or other constraints; CDW 700 encapsulates ML labeler 712, blind judgement human labeler 714, and open judgement human labeler 716, though it should be appreciated that a CDW can encapsulate any number of labelers of various types; each labeler instance may have a labeler instance score determined, for example, by QMS 750, and that corresponds to the probability that the labeler instance will produce an accurate label for a given task; the QMS 750 can score how often the labeler instance was correct when it labeled images as "tumor" and score how often the labeler instance was correct when it labeled images as "no tumor"; the answer-specific scores can be used in determining confidence estimates for an actual result output by the labeler instance; human labeler instance scores and ML labeler instance scores can be tied to specific labeling task types in the scoring system; for a labeler instance that can produce labels for multiple task types, QMS 750 may determine scores on the different task types as skill scores to differentiate the labeler instance's performance across different labeling task types; a labeler instance may have an associated response time (temporal cost) that is an estimate of how long it will take that labeler instance to perform a task; a labeler instance may have an associated price (monetary cost) that is an estimate of the price for that labeler instance to perform a task; for a particular labeling task, the workflow orchestrator 710 may select from labelers suited for that type of task; a labeler may also have a labeler score; a labeler's score corresponds to the probability that a labeler will produce an accurate label for a given task; a workflow orchestrator 710 to dynamically determine an execution path for processing a labeling request ( question) to produce a final labeled result based on confidence and cost constraint configuration; workflow orchestrator 710 uses the task information for the workflow task, the labeler descriptions, possibly dynamic labeler characteristics such as cost, availability, and timeliness, and quality metrics from QMS 750 to dynamically determine a path through the constituent labelers 712, 714, 716 to produce a labeled output for the task which satisfies a configured target confidence threshold (a minimum confidence threshold); the path may also be selected to minimize costs in one or more dimensions; workflow orchestrator 710 uses the labeler scores and costs associated with the labelers to determine one or more viable paths through the labelers; once a viable path is determined, workflow orchestrator 710 may route the labeling request through that path until a result reaches a threshold confidence target or the path is exhausted; the path may be re-evaluated and changed at any time based on the actual results produced by each consultation; the order of consultation of constituent labelers may proceed from least expensive labeler to most expensive labeler, for example, in an attempt to reach the confidence threshold target with the least cost; if ML labeler 712 is the least expensive labeler, blind judgement human labeler 714 is more expensive because labeling requests to blind judgement human labeler 714 are routed to human specialists who have a higher monetary cost based on compensation for work performed, and human labeler 716 is the most expensive of the constituent labelers because labeling requests to human labeler 716 are routed to human specialists with higher expertise and compensation levels than the human specialists associated with blind judgement human labeler 714, then workflow orchestrator 710 may favor consulting ML labeler 712 first, then human labeler 714, then human labeler 716; in the example of FIG. 8B, the labeling request is routed to ML labeler 712, which returns an answer for which a confidence estimate 806 is determined; the labeling request is then routed to blind judgement human labeler 714, which produces an answer for which a confidence estimate 808 is determined; as the target confidence threshold has not been reached, the labeling request can be routed to blind judgement human labeler 714 twice because there are labeler instances (human specialists 727) remaining in the pool of human labeler 714 that have not yet been consulted for the labeling request; the confidence estimate 810 for the answer produced by the second consultation with human labeler 714 may incorporate confidence estimates 806, 808, however, exceeds the target confidence threshold; thus, that answer can be used for the final labeled results of CDW 700; QMS 750 provides quality metrics that may be used to dynamically determine the execution path and determine if a label result received from a constituent labeler ( or agreed to by multiple labelers) exceeds the configured target confidence threshold; QMS 750 can continuously monitor and score labeler instances over time to generate, maintain, and improve confidence estimates; QMS 750 can identify the set of constraints on labelers required to achieve a desired confidence threshold for an overall result (which may be composed of results from multiple labelers); various methods of determining the complex granular confidence estimate may be used; e.g., approaches may be employed that take into account the actual value of the answer (e.g., precision versus recall considerations, which into account differences in false positives versus false negatives and different likelihoods of providing one wrong answer versus a different wrong answer when the true answer is x versus if it is y); logistic regression methods or other estimators may be applied to determine complex granular confidence estimates rather than closed form probability equations; an ML estimation model may be used for combining into a single probability value; logistic regression models can be trained for combining confidence estimates into a single probability value).
Claim 15
Prendki in view of Martin discloses all the elements as stated in Claim 14 and further discloses wherein computing labeling quality of each automatic labeling procedure in the set of the automatic labeling procedures comprises: computing the recall rate, the precision rate, the accuracy rate and a ratio of correctly identified negatives of each automatic labeling procedure in the set of the automatic labeling procedures respectively based on the result of labeling the quality inspection data by each of the automatic labeling procedures and the result of manual labeling the quality inspection data; obtaining labeling quality of each automatic labeling procedure in the set of the automatic labeling procedures through linear weighting of the recall rate, the precision rate, the accuracy rate and the ratio of correctly identified negatives (Martin, ¶¶ [0095]-[0147] with FIGS. 7 and 8A-D: ML labelers, human labelers and other labelers can be combined into a confidence-driven workflow (CDW); a CDW can thus be considered a labeler that encapsulates a collection of other labelers, and more particularly, a collection of labelers of the same arity; the encapsulated labelers can be consulted in sequence and their individual results incorporated into an overall result until a configured threshold confidence target for an overall result is reached; a labeling request may be received as a workflow task with accompanying task information, such as a task type description and constraints; examples of constraints include, but are not limited to, cost constraints in one or more dimensions (e.g., time limit, monetary limit), target threshold confidence or other constraints; CDW 700 encapsulates ML labeler 712, blind judgement human labeler 714, and open judgement human labeler 716, though it should be appreciated that a CDW can encapsulate any number of labelers of various types; each labeler instance may have a labeler instance score determined, for example, by QMS 750, and that corresponds to the probability that the labeler instance will produce an accurate label for a given task; the QMS 750 can score how often the labeler instance was correct when it labeled images as "tumor" and score how often the labeler instance was correct when it labeled images as "no tumor"; the answer-specific scores can be used in determining confidence estimates for an actual result output by the labeler instance; human labeler instance scores and ML labeler instance scores can be tied to specific labeling task types in the scoring system; for a labeler instance that can produce labels for multiple task types, QMS 750 may determine scores on the different task types as skill scores to differentiate the labeler instance's performance across different labeling task types; a labeler instance may have an associated response time (temporal cost) that is an estimate of how long it will take that labeler instance to perform a task; a labeler instance may have an associated price (monetary cost) that is an estimate of the price for that labeler instance to perform a task; for a particular labeling task, the workflow orchestrator 710 may select from labelers suited for that type of task; a labeler may also have a labeler score; a labeler's score corresponds to the probability that a labeler will produce an accurate label for a given task; a workflow orchestrator 710 to dynamically determine an execution path for processing a labeling request ( question) to produce a final labeled result based on confidence and cost constraint configuration; workflow orchestrator 710 uses the task information for the workflow task, the labeler descriptions, possibly dynamic labeler characteristics such as cost, availability, and timeliness, and quality metrics from QMS 750 to dynamically determine a path through the constituent labelers 712, 714, 716 to produce a labeled output for the task which satisfies a configured target confidence threshold (a minimum confidence threshold); the path may also be selected to minimize costs in one or more dimensions; workflow orchestrator 710 uses the labeler scores and costs associated with the labelers to determine one or more viable paths through the labelers; once a viable path is determined, workflow orchestrator 710 may route the labeling request through that path until a result reaches a threshold confidence target or the path is exhausted; the path may be re-evaluated and changed at any time based on the actual results produced by each consultation; the order of consultation of constituent labelers may proceed from least expensive labeler to most expensive labeler, for example, in an attempt to reach the confidence threshold target with the least cost; if ML labeler 712 is the least expensive labeler, blind judgement human labeler 714 is more expensive because labeling requests to blind judgement human labeler 714 are routed to human specialists who have a higher monetary cost based on compensation for work performed, and human labeler 716 is the most expensive of the constituent labelers because labeling requests to human labeler 716 are routed to human specialists with higher expertise and compensation levels than the human specialists associated with blind judgement human labeler 714, then workflow orchestrator 710 may favor consulting ML labeler 712 first, then human labeler 714, then human labeler 716; in the example of FIG. 8B, the labeling request is routed to ML labeler 712, which returns an answer for which a confidence estimate 806 is determined; the labeling request is then routed to blind judgement human labeler 714, which produces an answer for which a confidence estimate 808 is determined; as the target confidence threshold has not been reached, the labeling request can be routed to blind judgement human labeler 714 twice because there are labeler instances (human specialists 727) remaining in the pool of human labeler 714 that have not yet been consulted for the labeling request; the confidence estimate 810 for the answer produced by the second consultation with human labeler 714 may incorporate confidence estimates 806, 808, however, exceeds the target confidence threshold; thus, that answer can be used for the final labeled results of CDW 700; QMS 750 provides quality metrics that may be used to dynamically determine the execution path and determine if a label result received from a constituent labeler ( or agreed to by multiple labelers) exceeds the configured target confidence threshold; QMS 750 can continuously monitor and score labeler instances over time to generate, maintain, and improve confidence estimates; QMS 750 can identify the set of constraints on labelers required to achieve a desired confidence threshold for an overall result (which may be composed of results from multiple labelers); various methods of determining the complex granular confidence estimate may be used; e.g., approaches may be employed that take into account the actual value of the answer (e.g., precision versus recall considerations, which into account differences in false positives versus false negatives and different likelihoods of providing one wrong answer versus a different wrong answer when the true answer is x versus if it is y); logistic regression methods or other estimators may be applied to determine complex granular confidence estimates rather than closed form probability equations; an ML estimation model may be used for combining into a single probability value; logistic regression models can be trained for combining confidence estimates into a single probability value; ¶ [0164]: a score for a labeler instance can be configured as a weighted sum of scores calculated from different scoring action sources or configured to combine scoring actions of various types into a single scoring formula).
Claim 19
Prendki in view of Martin discloses all the elements as stated in Claim 3 and further discloses wherein determining whether the first automatic labeling procedure satisfies the first quality evaluation index comprises: extracting quality inspection data from the labeling data, the quality inspection data is a part of the labeling data; labeling the quality inspection data with the first automatic labeling procedure; determining whether the first automatic labeling procedure is greater than a quality control threshold of the first quality evaluation index based on a result of labeling the quality inspection data by the first automatic labeling procedure and the result of manual labeling the quality inspection data; if the first automatic labeling procedure is greater than the quality control threshold, determining that the first automatic labeling procedure satisfies the first quality evaluation index; if the first automatic labeling procedure is not greater than the quality control threshold, determining that the first automatic labeling procedure fails to satisfy the first quality evaluation index (Prendki, ¶¶ [0186]-[0189]: verify the quality of the labeling tasks completed by the labeling partners by evaluating the completed tasks; the recommendation computer system 306 can be programmed with an anomaly detection engine that creates a distribution of the annotations as returned by the labeling partner and looks for potential anomalous annotations such as out of range values; the anomaly detection engine can be programmed for generating a z-score distribution in various features of interest; recommendation computer system 306 can be programmed with an auto-labeling/pre-labeling process, using a pre-trained machine learning model to generate synthetic generation, which are checked against the manual annotations provided by the labeling partner; in ML-driven data curation, which is active learning-based; data is selected dynamically during a training process; the training process ends when the user reaches their budget or the remaining data contains no new relevant information; ¶¶ [0060], [0174], and [0226]-[0229] with FIGS. 1-2: at step 201, a customer uploads to the system labeling instructions, data to be labeled, and a model to be trained using the data; at step 202, the customer decides a labeling approach (e.g., ML-based approach, manual approach, or self-labeling approach); at step 203, the curation process begins, and at step 204, the customer's model and data are sent to the data curation engine; the system's data preparation and data curation services employ an Active Learning approach where portions of the data are prioritized so the portions that are most effective in training the model are labeled first; the curation engine analyzes the data and identifies a first batch of data that it believes will be effective in training the data; the first batch of data is provided to a labeling partner in a similar fashion as the processes described in FIG. 1; once the labeling partner completes labeling the first batch of data, the labeled data is used to train the customer's model, as described in step 206; then, the model's performance is evaluated; at steps 207 and 208, if the system determines that the model's performance is acceptable, the system deems the model as sufficiently trained and provides the trained model and labeled data back to the customer (i.e., the experiment ends); the experiment may also end if the system determines that the customer's budget has ran out or if the model has reached its optimal performance; at step 205, if the system determines that the model's performance is unacceptable, the next loop of the experiment begins; going back to step 204, the data curation engine identifies the next batch of data and provides this batch to the labeling partner along with all of the previous batches of data; once the labeling partner completes labeling a batch of the data, the experiment continues until the model's performance reaches an acceptable level, the budget is reached, or optimal performance has been achieved because the remaining data is determined as redundant or irrelevant for training the model; ¶¶ [0203]-[0204]: provide quality assurances by identifying problematic labels; employ an adversarial process that involves providing deceptive or faulty data and evaluating the response of a service provider to receiving the deceptive or faulty data; provide the data labeled by one labeling partner to another labeling partner for cross-validation) (Martin, ¶¶ [0095]-[0147] with FIGS. 7 and 8A-D: ML labelers, human labelers and other labelers can be combined into a confidence-driven workflow (CDW); a CDW can thus be considered a labeler that encapsulates a collection of other labelers, and more particularly, a collection of labelers of the same arity; the encapsulated labelers can be consulted in sequence and their individual results incorporated into an overall result until a configured threshold confidence target for an overall result is reached; a labeling request may be received as a workflow task with accompanying task information, such as a task type description and constraints; examples of constraints include, but are not limited to, cost constraints in one or more dimensions (e.g., time limit, monetary limit), target threshold confidence or other constraints; CDW 700 encapsulates ML labeler 712, blind judgement human labeler 714, and open judgement human labeler 716, though it should be appreciated that a CDW can encapsulate any number of labelers of various types; each labeler instance may have a labeler instance score determined, for example, by QMS 750, and that corresponds to the probability that the labeler instance will produce an accurate label for a given task; the QMS 750 can score how often the labeler instance was correct when it labeled images as "tumor" and score how often the labeler instance was correct when it labeled images as "no tumor"; the answer-specific scores can be used in determining confidence estimates for an actual result output by the labeler instance; human labeler instance scores and ML labeler instance scores can be tied to specific labeling task types in the scoring system; for a labeler instance that can produce labels for multiple task types, QMS 750 may determine scores on the different task types as skill scores to differentiate the labeler instance's performance across different labeling task types; a labeler instance may have an associated response time (temporal cost) that is an estimate of how long it will take that labeler instance to perform a task; a labeler instance may have an associated price (monetary cost) that is an estimate of the price for that labeler instance to perform a task; for a particular labeling task, the workflow orchestrator 710 may select from labelers suited for that type of task; a labeler may also have a labeler score; a labeler's score corresponds to the probability that a labeler will produce an accurate label for a given task; a workflow orchestrator 710 to dynamically determine an execution path for processing a labeling request ( question) to produce a final labeled result based on confidence and cost constraint configuration; workflow orchestrator 710 uses the task information for the workflow task, the labeler descriptions, possibly dynamic labeler characteristics such as cost, availability, and timeliness, and quality metrics from QMS 750 to dynamically determine a path through the constituent labelers 712, 714, 716 to produce a labeled output for the task which satisfies a configured target confidence threshold (a minimum confidence threshold); the path may also be selected to minimize costs in one or more dimensions; workflow orchestrator 710 uses the labeler scores and costs associated with the labelers to determine one or more viable paths through the labelers; once a viable path is determined, workflow orchestrator 710 may route the labeling request through that path until a result reaches a threshold confidence target or the path is exhausted; the path may be re-evaluated and changed at any time based on the actual results produced by each consultation; the order of consultation of constituent labelers may proceed from least expensive labeler to most expensive labeler, for example, in an attempt to reach the confidence threshold target with the least cost; if ML labeler 712 is the least expensive labeler, blind judgement human labeler 714 is more expensive because labeling requests to blind judgement human labeler 714 are routed to human specialists who have a higher monetary cost based on compensation for work performed, and human labeler 716 is the most expensive of the constituent labelers because labeling requests to human labeler 716 are routed to human specialists with higher expertise and compensation levels than the human specialists associated with blind judgement human labeler 714, then workflow orchestrator 710 may favor consulting ML labeler 712 first, then human labeler 714, then human labeler 716; in the example of FIG. 8B, the labeling request is routed to ML labeler 712, which returns an answer for which a confidence estimate 806 is determined; the labeling request is then routed to blind judgement human labeler 714, which produces an answer for which a confidence estimate 808 is determined; as the target confidence threshold has not been reached, the labeling request can be routed to blind judgement human labeler 714 twice because there are labeler instances (human specialists 727) remaining in the pool of human labeler 714 that have not yet been consulted for the labeling request; the confidence estimate 810 for the answer produced by the second consultation with human labeler 714 may incorporate confidence estimates 806, 808, however, exceeds the target confidence threshold; thus, that answer can be used for the final labeled results of CDW 700; QMS 750 provides quality metrics that may be used to dynamically determine the execution path and determine if a label result received from a constituent labeler ( or agreed to by multiple labelers) exceeds the configured target confidence threshold; QMS 750 can continuously monitor and score labeler instances over time to generate, maintain, and improve confidence estimates; QMS 750 can identify the set of constraints on labelers required to achieve a desired confidence threshold for an overall result (which may be composed of results from multiple labelers); various methods of determining the complex granular confidence estimate may be used; e.g., approaches may be employed that take into account the actual value of the answer (e.g., precision versus recall considerations, which into account differences in false positives versus false negatives and different likelihoods of providing one wrong answer versus a different wrong answer when the true answer is x versus if it is y); logistic regression methods or other estimators may be applied to determine complex granular confidence estimates rather than closed form probability equations; an ML estimation model may be used for combining into a single probability value; logistic regression models can be trained for combining confidence estimates into a single probability value).
Claims 10-11 are rejected under 35 U.S.C. 103 as being unpatentable over Prendki in view of Martin as applied to Claim 2 above, and further in view of ZOU (CN 111444166 A, pub. date: 07/24/2020), hereinafter ZOU.
Claim 10
Prendki in view of Martin discloses all the elements as stated in Claim 2 and further discloses wherein determining whether the existing labeling procedure satisfies the first quality evaluation index comprises: extracting a preset quality inspection (Prendki, ¶¶ [0186]-[0189]: verify the quality of the labeling tasks completed by the labeling partners by evaluating the completed tasks; the recommendation computer system 306 can be programmed with an anomaly detection engine that creates a distribution of the annotations as returned by the labeling partner and looks for potential anomalous annotations such as out of range values; the anomaly detection engine can be programmed for generating a z-score distribution in various features of interest; recommendation computer system 306 can be programmed with an auto-labeling/pre-labeling process, using a pre-trained machine learning model to generate synthetic generation, which are checked against the manual annotations provided by the labeling partner; in ML-driven data curation, which is active learning-based; data is selected dynamically during a training process; the training process ends when the user reaches their budget or the remaining data contains no new relevant information; ¶¶ [0060], [0174], and [0226]-[0229] with FIGS. 1-2: at step 201, a customer uploads to the system labeling instructions, data to be labeled, and a model to be trained using the data; at step 202, the customer decides a labeling approach (e.g., ML-based approach, manual approach, or self-labeling approach); at step 203, the curation process begins, and at step 204, the customer's model and data are sent to the data curation engine; the system's data preparation and data curation services employ an Active Learning approach where portions of the data are prioritized so the portions that are most effective in training the model are labeled first; the curation engine analyzes the data and identifies a first batch of data that it believes will be effective in training the data; the first batch of data is provided to a labeling partner in a similar fashion as the processes described in FIG. 1; once the labeling partner completes labeling the first batch of data, the labeled data is used to train the customer's model, as described in step 206; then, the model's performance is evaluated; at steps 207 and 208, if the system determines that the model's performance is acceptable, the system deems the model as sufficiently trained and provides the trained model and labeled data back to the customer (i.e., the experiment ends); the experiment may also end if the system determines that the customer's budget has ran out or if the model has reached its optimal performance; at step 205, if the system determines that the model's performance is unacceptable, the next loop of the experiment begins; going back to step 204, the data curation engine identifies the next batch of data and provides this batch to the labeling partner along with all of the previous batches of data; once the labeling partner completes labeling a batch of the data, the experiment continues until the model's performance reaches an acceptable level, the budget is reached, or optimal performance has been achieved because the remaining data is determined as redundant or irrelevant for training the model; ¶¶ [0203]-[0204]: provide quality assurances by identifying problematic labels; employ an adversarial process that involves providing deceptive or faulty data and evaluating the response of a service provider to receiving the deceptive or faulty data; provide the data labeled by one labeling partner to another labeling partner for cross-validation) (Martin, ¶¶ [0095]-[0147] with FIGS. 7 and 8A-D: ML labelers, human labelers and other labelers can be combined into a confidence-driven workflow (CDW); a CDW can thus be considered a labeler that encapsulates a collection of other labelers, and more particularly, a collection of labelers of the same arity; the encapsulated labelers can be consulted in sequence and their individual results incorporated into an overall result until a configured threshold confidence target for an overall result is reached; a labeling request may be received as a workflow task with accompanying task information, such as a task type description and constraints; examples of constraints include, but are not limited to, cost constraints in one or more dimensions (e.g., time limit, monetary limit), target threshold confidence or other constraints; CDW 700 encapsulates ML labeler 712, blind judgement human labeler 714, and open judgement human labeler 716, though it should be appreciated that a CDW can encapsulate any number of labelers of various types; each labeler instance may have a labeler instance score determined, for example, by QMS 750, and that corresponds to the probability that the labeler instance will produce an accurate label for a given task; the QMS 750 can score how often the labeler instance was correct when it labeled images as "tumor" and score how often the labeler instance was correct when it labeled images as "no tumor"; the answer-specific scores can be used in determining confidence estimates for an actual result output by the labeler instance; human labeler instance scores and ML labeler instance scores can be tied to specific labeling task types in the scoring system; for a labeler instance that can produce labels for multiple task types, QMS 750 may determine scores on the different task types as skill scores to differentiate the labeler instance's performance across different labeling task types; a labeler instance may have an associated response time (temporal cost) that is an estimate of how long it will take that labeler instance to perform a task; a labeler instance may have an associated price (monetary cost) that is an estimate of the price for that labeler instance to perform a task; for a particular labeling task, the workflow orchestrator 710 may select from labelers suited for that type of task; a labeler may also have a labeler score; a labeler's score corresponds to the probability that a labeler will produce an accurate label for a given task; a workflow orchestrator 710 to dynamically determine an execution path for processing a labeling request ( question) to produce a final labeled result based on confidence and cost constraint configuration; workflow orchestrator 710 uses the task information for the workflow task, the labeler descriptions, possibly dynamic labeler characteristics such as cost, availability, and timeliness, and quality metrics from QMS 750 to dynamically determine a path through the constituent labelers 712, 714, 716 to produce a labeled output for the task which satisfies a configured target confidence threshold (a minimum confidence threshold); the path may also be selected to minimize costs in one or more dimensions; workflow orchestrator 710 uses the labeler scores and costs associated with the labelers to determine one or more viable paths through the labelers; once a viable path is determined, workflow orchestrator 710 may route the labeling request through that path until a result reaches a threshold confidence target or the path is exhausted; the path may be re-evaluated and changed at any time based on the actual results produced by each consultation; the order of consultation of constituent labelers may proceed from least expensive labeler to most expensive labeler, for example, in an attempt to reach the confidence threshold target with the least cost; if ML labeler 712 is the least expensive labeler, blind judgement human labeler 714 is more expensive because labeling requests to blind judgement human labeler 714 are routed to human specialists who have a higher monetary cost based on compensation for work performed, and human labeler 716 is the most expensive of the constituent labelers because labeling requests to human labeler 716 are routed to human specialists with higher expertise and compensation levels than the human specialists associated with blind judgement human labeler 714, then workflow orchestrator 710 may favor consulting ML labeler 712 first, then human labeler 714, then human labeler 716; in the example of FIG. 8B, the labeling request is routed to ML labeler 712, which returns an answer for which a confidence estimate 806 is determined; the labeling request is then routed to blind judgement human labeler 714, which produces an answer for which a confidence estimate 808 is determined; as the target confidence threshold has not been reached, the labeling request can be routed to blind judgement human labeler 714 twice because there are labeler instances (human specialists 727) remaining in the pool of human labeler 714 that have not yet been consulted for the labeling request; the confidence estimate 810 for the answer produced by the second consultation with human labeler 714 may incorporate confidence estimates 806, 808, however, exceeds the target confidence threshold; thus, that answer can be used for the final labeled results of CDW 700; QMS 750 provides quality metrics that may be used to dynamically determine the execution path and determine if a label result received from a constituent labeler ( or agreed to by multiple labelers) exceeds the configured target confidence threshold; QMS 750 can continuously monitor and score labeler instances over time to generate, maintain, and improve confidence estimates; QMS 750 can identify the set of constraints on labelers required to achieve a desired confidence threshold for an overall result (which may be composed of results from multiple labelers); various methods of determining the complex granular confidence estimate may be used; e.g., approaches may be employed that take into account the actual value of the answer (e.g., precision versus recall considerations, which into account differences in false positives versus false negatives and different likelihoods of providing one wrong answer versus a different wrong answer when the true answer is x versus if it is y); logistic regression methods or other estimators may be applied to determine complex granular confidence estimates rather than closed form probability equations; an ML estimation model may be used for combining into a single probability value; logistic regression models can be trained for combining confidence estimates into a single probability value).
Prendki in view of Martin fails to explicitly disclose extracting a preset quality inspection ratio of data from the automatic labeling data and setting the preset quality inspection ratio of data as quality inspection data.
ZOU teaches a system and method relating to labeling data (ZOU, Abstract), wherein extracting a preset quality inspection ratio of data from the automatic labeling data and setting the preset quality inspection ratio of data as quality inspection data (ZOU, ¶¶ [0004]-[0014] and [0022]-[0038] with FIG. 1: provide an automatic quality inspection method for labeled data, the method comprises: S1, obtain the data to be labeled, and divide the data to be labeled into m batches, each batch containing m data items; S2, extract a preset number of data from each batch of data for labeling, as the initial labeled standard dataset; S3, add the initial standard dataset to each batch of data, and label the data in each batch that contains the initial standard dataset; S4, by detecting the data labeled in step S3, the background automatically calculates the accuracy of the initial standard dataset; S5. determine whether the accuracy rate has reached the preset standard value; if yes, pass the automatic quality inspection; otherwise, execute step S2 to re-label; preferably, in step S2, the preset quantity is defined as m1, which satisfies m1 = 10% * m)
Prendki in view of Martin, and ZOU are analogous art because they are from the same field of endeavor, a system and method relating to labeling data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of ZOU to Prendki in view of Martin. Motivation for doing so would save time and effort for ensuring quality of labeled data (.
Claim 11
Prendki in view of Martin and ZOU discloses all the elements as stated in Claim 10 and further discloses wherein the first quality evaluation index comprises a recall rate, a precision rate, an accuracy rate, a false positive rate and a false negative rate; and wherein determining whether the existing automatic labeling procedure satisfies the first quality evaluation index based on the result of labeling the quality inspection data by the existing automatic labeling procedure and the result of manual labeling the quality inspection data comprises: computing the recall rate, the precision rate, the accuracy rate, the false positive rate and the false negative rate based on the result of labeling the quality inspection data by the existing automatic labeling procedure and the result of manual labeling the quality inspection data; determining whether the existing automatic labeling procedure satisfies the first quality evaluation index based on whether the recall rate, the precision rate, the accuracy rate, the false positive rate and the false negative rate are respectively greater than a recall rate threshold, a precision rate threshold, an accuracy rate threshold, a false positive rate threshold and a false negative rate threshold; in case that the recall rate, the precision rate, the accuracy rate, the false positive rate and the false negative rate are respectively greater than the recall rate threshold, the precision rate threshold, the accuracy rate threshold, the false positive rate threshold and the false negative rate threshold, determining that the existing automatic labeling procedure satisfies the first quality evaluation index; in case that at least one of the recall rate, the precision rate, the accuracy rate, the false positive rate and the false negative rate is not greater than the recall rate threshold, the precision rate threshold, the accuracy rate threshold, the false positive rate threshold and the false negative rate threshold, determining that the existing automatic labeling procedure fails to satisfy the first quality evaluation index (Martin, ¶¶ [0095]-[0147] with FIGS. 7 and 8A-D: ML labelers, human labelers and other labelers can be combined into a confidence-driven workflow (CDW); a CDW can thus be considered a labeler that encapsulates a collection of other labelers, and more particularly, a collection of labelers of the same arity; the encapsulated labelers can be consulted in sequence and their individual results incorporated into an overall result until a configured threshold confidence target for an overall result is reached; a labeling request may be received as a workflow task with accompanying task information, such as a task type description and constraints; examples of constraints include, but are not limited to, cost constraints in one or more dimensions (e.g., time limit, monetary limit), target threshold confidence or other constraints; CDW 700 encapsulates ML labeler 712, blind judgement human labeler 714, and open judgement human labeler 716, though it should be appreciated that a CDW can encapsulate any number of labelers of various types; each labeler instance may have a labeler instance score determined, for example, by QMS 750, and that corresponds to the probability that the labeler instance will produce an accurate label for a given task; the QMS 750 can score how often the labeler instance was correct when it labeled images as "tumor" and score how often the labeler instance was correct when it labeled images as "no tumor"; the answer-specific scores can be used in determining confidence estimates for an actual result output by the labeler instance; human labeler instance scores and ML labeler instance scores can be tied to specific labeling task types in the scoring system; for a labeler instance that can produce labels for multiple task types, QMS 750 may determine scores on the different task types as skill scores to differentiate the labeler instance's performance across different labeling task types; a labeler instance may have an associated response time (temporal cost) that is an estimate of how long it will take that labeler instance to perform a task; a labeler instance may have an associated price (monetary cost) that is an estimate of the price for that labeler instance to perform a task; for a particular labeling task, the workflow orchestrator 710 may select from labelers suited for that type of task; a labeler may also have a labeler score; a labeler's score corresponds to the probability that a labeler will produce an accurate label for a given task; a workflow orchestrator 710 to dynamically determine an execution path for processing a labeling request ( question) to produce a final labeled result based on confidence and cost constraint configuration; workflow orchestrator 710 uses the task information for the workflow task, the labeler descriptions, possibly dynamic labeler characteristics such as cost, availability, and timeliness, and quality metrics from QMS 750 to dynamically determine a path through the constituent labelers 712, 714, 716 to produce a labeled output for the task which satisfies a configured target confidence threshold (a minimum confidence threshold); the path may also be selected to minimize costs in one or more dimensions; workflow orchestrator 710 uses the labeler scores and costs associated with the labelers to determine one or more viable paths through the labelers; once a viable path is determined, workflow orchestrator 710 may route the labeling request through that path until a result reaches a threshold confidence target or the path is exhausted; the path may be re-evaluated and changed at any time based on the actual results produced by each consultation; the order of consultation of constituent labelers may proceed from least expensive labeler to most expensive labeler, for example, in an attempt to reach the confidence threshold target with the least cost; if ML labeler 712 is the least expensive labeler, blind judgement human labeler 714 is more expensive because labeling requests to blind judgement human labeler 714 are routed to human specialists who have a higher monetary cost based on compensation for work performed, and human labeler 716 is the most expensive of the constituent labelers because labeling requests to human labeler 716 are routed to human specialists with higher expertise and compensation levels than the human specialists associated with blind judgement human labeler 714, then workflow orchestrator 710 may favor consulting ML labeler 712 first, then human labeler 714, then human labeler 716; in the example of FIG. 8B, the labeling request is routed to ML labeler 712, which returns an answer for which a confidence estimate 806 is determined; the labeling request is then routed to blind judgement human labeler 714, which produces an answer for which a confidence estimate 808 is determined; as the target confidence threshold has not been reached, the labeling request can be routed to blind judgement human labeler 714 twice because there are labeler instances (human specialists 727) remaining in the pool of human labeler 714 that have not yet been consulted for the labeling request; the confidence estimate 810 for the answer produced by the second consultation with human labeler 714 may incorporate confidence estimates 806, 808, however, exceeds the target confidence threshold; thus, that answer can be used for the final labeled results of CDW 700; QMS 750 provides quality metrics that may be used to dynamically determine the execution path and determine if a label result received from a constituent labeler ( or agreed to by multiple labelers) exceeds the configured target confidence threshold; QMS 750 can continuously monitor and score labeler instances over time to generate, maintain, and improve confidence estimates; QMS 750 can identify the set of constraints on labelers required to achieve a desired confidence threshold for an overall result (which may be composed of results from multiple labelers); various methods of determining the complex granular confidence estimate may be used; e.g., approaches may be employed that take into account the actual value of the answer (e.g., precision versus recall considerations, which into account differences in false positives versus false negatives and different likelihoods of providing one wrong answer versus a different wrong answer when the true answer is x versus if it is y); logistic regression methods or other estimators may be applied to determine complex granular confidence estimates rather than closed form probability equations; an ML estimation model may be used for combining into a single probability value; logistic regression models can be trained for combining confidence estimates into a single probability value; ¶ [0164]: a score for a labeler instance can be configured as a weighted sum of scores calculated from different scoring action sources or configured to combine scoring actions of various types into a single scoring formula).
Claims 16-18 are rejected under 35 U.S.C. 103 as being unpatentable over Prendki in view of Martin as applied to Claim 3 above, and further in view of Li et al. ("How to Simplify Search: Classification-wise Pareto Evolution for One-shot Neural Architecture Search", arXiv:2109.07582v1, Sep 14, 2021, pp. 1-6), hereinafter Li.
Claim 16
Prendki in view of Martin discloses all the elements as stated in Claim 3 and further discloses wherein determining the first automatic labeling procedure based on the labeling quality and the labeling metric comprises: computing a (Prendki, ¶ [0050]: efficiently compare the capabilities, availability, performance, and costs of various labeling providers; ¶ [0057]:recommendation computer system 306 is programmed using the real-time evaluation instructions 308 to query the provider metrics database 310 and to select, automatically and based upon programmed recommendation or evaluation algorithms, one or more of the label service providers 314, 316, 318 that is optimal to perform one or more labeling tasks on the user dataset; ¶¶ [0114]-[0154]: collect various information of the labeling partners, which may be used by the optimization system to find the most suitable labeling partner for a customer-requested task; examples of such information that the system collects of the labeling partners may include the following: list of the types of tasks that the partner can deliver on (depends on tooling and labor workforce); available clearance; annotator information (e.g., type of tasks annotators are capable of working on); historical performance of the annotators (e.g., speed and accuracy); desired workload of annotators; and real-time availability of annotators; compute and maintain various metrics of the labeling providers; examples of such metrics may include a) average labeling quality (e.g., accuracy) for various types of tasks; b) average labeling speed for various types of tasks; c) typical availability; d) average price, and deviation from standard price; e) time since last offered task; f) time of idle (e.g., no assigned task); g) time taken to resolve labeling issues identified by the quality control process; and number of iterations required to solve a labeling issue; h) response time to the system's inquiry/notifications; and i) acceptance/rejection/ignore statistics of labeling tasks; utilize a multi-objective optimization process to find a labeling partner that is most likely to produce the highest quality of work; examples of the factors considered during the optimization process include, but are not limited to, the following: a) turn-around time (e.g., on average how long does the labeling partner need to complete and return a task); b) labeling quality (e.g., how accurate is the labeled data); c) average cost (e.g., the average cost charged by the labeling partner for various types of tasks); d) revenue (e.g., how much did the labeling partner earn for a given period of time); e) number of, and type of, tasks previously offered to the labeling partner; f) number of labeling partners within the computer-implemented labeling marketplace; and g) customer's preferences/criteria; the optimization process may be programmed using any of: the Dantzig simplex algorithm, extensions or variants thereof; combinatorial algorithms; or quantum optimization algorithms; Claim 10: the identifying by multi-objective optimization based on: an average time required by the labeling partners to complete and return a task; an average labeling accuracy of the labeling partners; an average price of the labeling partners; an average revenue of the labeling partners over a given period of time; a number of tasks previously offered to the labeling partners; and types of tasks previously offered to the labeling partners; ¶¶ [0174] and [0228]-[0229]: the system may end the experiment if it determines that the model has reached its optimal performance; the system determines that the model has reached its optimal performance if the portions of the data identified by the data curation engine as being effective at training the model have all been labeled and only the portions that are identified as redundant or irrelevant in training the model are left; ¶ [0206]: provides universal labeling instructions that are optimized for all labeling partners within the computer-implemented labeling marketplace; the universal labeling instructions are designed to be universally compatible with all of the labeling partners within the computer-implemented labeling marketplace by considering each of the labeling partners' capabilities at both the provider-level and annotator-level and the tooling available to the partners; ¶ [0220]: applies an optimization process that evaluates real-time and historical metrics of all the labeling partners remaining in the selection pool; ¶ [0214]: with Ranking-based Active Learning, the system is programmed to re-rank data; rather than predicting which record should be labeled in the context of a specific loop, the system is programmed to continuously re-rank records in the dataset based on qualitative aspects of the records (e.g., how effective the dataset is in training the model), which allows the system to select data not only for loop n, but the subsequent ones most effective in training the model as well) (Martin, ¶ [0029]: after a step in the path, the labeling platform will have more information about the actual confidence and costs so far in the execution path and a redetermination of candidate paths can be performed help optimize the path from the current point in the execution path forward, whether or not the expectations for the current point in the path have been met; ¶ [0041]: the CDW dynamically determines a cost-optimized execution path to meet a target confidence threshold for a labeling request; ¶ [0118]: CDW can be configured dynamically determine an execution path to optimize the confidence of the overall labeling result that incorporates the individual labeling results obtained so far for the labeling request; ¶ [0186]: if there are cost constraints (temporal, monetary or other cost) and/or confidence constraints, workflow orchestrator 710 performs a path search to find the optimal set of labelers (step 1114); ¶¶ [0197]-[0204]: workflow orchestrator 710 can pick the optimal path from the possibly multiple solutions found that meet the constraints; workflow orchestrator can optimize the path according to one or more of: 1) Cost optimized; 2) Time optimized; 3) Confidence optimized; combinations of these can be used by optimizing a set of confidence constraints for a range, then optimizing the subset for a different range on a different variable).
Prendki in view of Martin fails to explicitly disclose the Pareto dominance relation for performing optimization.
Li teaches a system and a method for performing optimization (Li, Abstract and 1st paragraph of Section I in Page 1), wherein the Pareto dominance relation for performing optimization (Li, Abstract in Page 1: in the deployment of deep neural models, how to effectively and automatically find feasible deep models under diverse design objectives is fundamental. Most existing neural architecture search (NAS) methods utilize surrogates to predict the detailed performance (e.g., accuracy and model size) of a candidate architecture during the search, which however is complicated and inefficient; in contrast, we aim to learn an efficient Pareto classifier to simplify the search process of NAS by transforming the complex multi-objective NAS task into a simple Pareto-dominance classification task; propose a classification wise Pareto evolution approach for one-shot NAS, where an online classifier is trained to predict the dominance relationship between the candidate and constructed reference architectures, instead of using surrogates to fit the objective functions; the main contribution of this study is to change supernet adaption into a Pareto classifier; besides, we design two adaptive schemes to select the reference set of architectures for constructing classification boundary and regulate the rate of positive samples over negative ones, respectively; Section I of Pages 1-2: NEURAL architecture search (NAS) is shown to be promising for the automatic design of task-specific deep neural networks (DNNs), instead of the traditional manual design based on extensive human expertise [1], [2], [3], [4]; consequently, NAS has received a surge of attention from the community of deep learning, largely owing to its superiority in optimizing the architecture and weights of DNN [2]; in other words, NAS can obtain desirable architectures as human experts do, or even far more innovative architectures [3]; although EA-based NAS is gradient free and insensitive to the complexity of the objectives [3], it suffers complicated and costly evaluation of candidate architectures on specific tasks; another issue of NAS is how to effectively optimize multiple conflicting design objectives (e.g., accuracy and model size) in solving practical applications; although many multi-objective selection strategies (e.g., the non-dominate sorting strategy [18]) have been integrated into NAS to drive the search towards the Pareto front over different objectives [19], there is still a big gap to improve the effectiveness [20]; specifically, if the objectives are not correctly evaluated, the Pareto dominance relationships between architectures will be easily misjudged, and the subsequent search process of NAS will be misled [21]; specifically, the classifier aims to predict the dominance relationship of each candidate architecture over a set of reference points (or architectures), and classify the candidate architectures into two (good and poor) classes; the samples in the good class are the points which either dominate or are non-dominant with the reference points; the poor class contains the samples dominated by the reference points; in this way, the complex architecture evaluation task is transformed into a Pareto-dominance classification task; besides, we design an adaptive clustering-based selection method and an α-dominance evaluation strategy to produce effective samples to avoid the class imbalance problem that the number of samples in one class is far less or more than that of another class; hence, high accuracy of the classification will be achieved; Section 3 with FIG. 1 and Algorithms 1-7 in Pages 3-8: Fig. 1 shows the framework of CENAS, which includes four main components: a multi-objective search routine, an ensemble Pareto classifier, a set of reference points, and an adaptive supernet; especially, the Pareto classifier is used to divide offspring architectures into two (good and poor) classes, while the set of reference points are used to construct training samples for classifier; an initial population of N architectures (subnets) is generated via the uniform sampling from the search space; then, the performance objectives (f1,…,fm) of each architecture in P are evaluated using weights inherited from the supernet on the validation set; a set of reference points SR is required to be In this process, an angle-based cluster strategy is used to make the reference points adapt to the shape of the Pareto front; moreover, a relaxed dominance strategy is employed to adjust the rate of positive and negative samples, so as to avoid the class imbalance problem selected from Arc to construct the classification boundary; the architectures in the archive are categorized into two classes (class1-good/class2-poor) according to the classification boundary (see Algorithm 3); then these samples are divided into the training data set and test data set for the training of a Pareto classifier, which is based on the ensemble of support vector machines (SVMs) (see Algorithm 4 and Algorithm 5); reproduction operators like crossover and mutation are performed on the parent solutions to create offspring solutions; promising architectures from the offspring are selected by the proposed classifier as the parent individuals of the next generation (see Algorithm 6); the archive Arc aims to keep track of promising architectures gained from the evolutionary search process so far, and it is updated using the obtained population P; then, the task-specific supernet is trained using the subnets sampled from architectures in the archive (see Algorithm 7); from the archive, we can obtain desired DNNs having the best trade-off between different objectives; present a densely-connected-like search space as shown in Fig. 2 (top left); specifically, a series of cells, i.e., the building blocks of an architecture, are connected in a feed-forward manner; two types of cells are designed: the normal cells aim to preserve the input size of data (or feature map), and the reduction cells to reduce the input spatial size; our goal is to find a suitable routing-like path between these cells to form a final architecture; in order to train a classifier to distinguish the candidate architectures, we select a set of reference points (architectures) as classification boundary to divide the samples into two (i.e., good and poor) classes; Apparently, the selected reference points determine the labeling of samples; Fig. 3 gives an example to show the impact of the position of reference points on the classification; design an adaptive clustering-based approach to select reference points from the obtained architectures. The main idea is to use clustering to help detect the distribution of the obtained solutions and guide the selection of reference points in line with the Pareto front; the details are given in Algorithm 2; non-domination sorting (Line 1): the nondominated fronts (PF1, PF2, …) of solutions are obtained by conducting conventional non-dominated sorting on the population from scratch [18], where the solutions in each front are nondominated with each other. The nondominated sorting is detailed in Section S1 of the supplementary material; to alleviate the computation cost of evaluating subnets under multiple objectives, we build an online Pareto classifier, which is to directly predict the Pareto-dominance classification of a subnet without training; Pareto dominance comparison is used to construct the labels of samples in the following way: if a solution is dominated by the set of reference points, then it will be labeled as a negative sample; otherwise it is a positive one; in the labeling process, the use of original Pareto-dominance may result in a class imbalance phenomenon; to handle this issue, we use α-domination criterion [37] to re-label samples (Lines 13-28 in Algorithm 3); the α-domination criterion. The α-domination is a relaxed Pareto-dominance by introducing a parameter α to control the dominated region; In Fig. 4, it can be seen that the α-domination is able to regulate the dominated region of solutions, and control the number of dominated solutions).
Prendki in view of Martin, and LI are analogous art because they are from the same field of endeavor, a system and method relating to labeling data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention to apply the teaching of Li to Prendki in view of Martin. Motivation for doing so would (1) s.
Claim 17
Prendki in view of Martin and Li discloses all the elements as stated in Claim 16 and further discloses wherein computing Pareto dominance relation between any two automatic labeling procedures in the set of the automatic labeling procedures based on the labeling quality and the labeling metric comprises: establishing a set of labeling qualities and a set of labeling metrics corresponding to the set of automatic labeling procedures; obtaining a first labeling metric of a first automatic labeling procedure in the set of automatic labeling procedures; determining whether the first labeling metric is smaller than or equal to a labeling metric threshold; if the first labeling metric is smaller than the labeling metric threshold, traversing the set of labeling qualities and the set labeling metrics to calculate Pareto dominance relation between any two automatic labeling procedures in the set of the automatic labeling procedures (Prendki, ¶ [0050]: efficiently compare the capabilities, availability, performance, and costs of various labeling providers; ¶ [0057]:recommendation computer system 306 is programmed using the real-time evaluation instructions 308 to query the provider metrics database 310 and to select, automatically and based upon programmed recommendation or evaluation algorithms, one or more of the label service providers 314, 316, 318 that is optimal to perform one or more labeling tasks on the user dataset; ¶¶ [0114]-[0154]: collect various information of the labeling partners, which may be used by the optimization system to find the most suitable labeling partner for a customer-requested task; examples of such information that the system collects of the labeling partners may include the following: list of the types of tasks that the partner can deliver on (depends on tooling and labor workforce); available clearance; annotator information (e.g., type of tasks annotators are capable of working on); historical performance of the annotators (e.g., speed and accuracy); desired workload of annotators; and real-time availability of annotators; compute and maintain various metrics of the labeling providers; examples of such metrics may include a) average labeling quality (e.g., accuracy) for various types of tasks; b) average labeling speed for various types of tasks; c) typical availability; d) average price, and deviation from standard price; e) time since last offered task; f) time of idle (e.g., no assigned task); g) time taken to resolve labeling issues identified by the quality control process; and number of iterations required to solve a labeling issue; h) response time to the system's inquiry/notifications; and i) acceptance/rejection/ignore statistics of labeling tasks; utilize a multi-objective optimization process to find a labeling partner that is most likely to produce the highest quality of work; examples of the factors considered during the optimization process include, but are not limited to, the following: a) turn-around time (e.g., on average how long does the labeling partner need to complete and return a task); b) labeling quality (e.g., how accurate is the labeled data); c) average cost (e.g., the average cost charged by the labeling partner for various types of tasks); d) revenue (e.g., how much did the labeling partner earn for a given period of time); e) number of, and type of, tasks previously offered to the labeling partner; f) number of labeling partners within the computer-implemented labeling marketplace; and g) customer's preferences/criteria; the optimization process may be programmed using any of: the Dantzig simplex algorithm, extensions or variants thereof; combinatorial algorithms; or quantum optimization algorithms; Claim 10: the identifying by multi-objective optimization based on: an average time required by the labeling partners to complete and return a task; an average labeling accuracy of the labeling partners; an average price of the labeling partners; an average revenue of the labeling partners over a given period of time; a number of tasks previously offered to the labeling partners; and types of tasks previously offered to the labeling partners; ¶¶ [0174] and [0228]-[0229]: the system may end the experiment if it determines that the model has reached its optimal performance; the system determines that the model has reached its optimal performance if the portions of the data identified by the data curation engine as being effective at training the model have all been labeled and only the portions that are identified as redundant or irrelevant in training the model are left; ¶ [0206]: provides universal labeling instructions that are optimized for all labeling partners within the computer-implemented labeling marketplace; the universal labeling instructions are designed to be universally compatible with all of the labeling partners within the computer-implemented labeling marketplace by considering each of the labeling partners' capabilities at both the provider-level and annotator-level and the tooling available to the partners; ¶ [0220]: applies an optimization process that evaluates real-time and historical metrics of all the labeling partners remaining in the selection pool; ¶ [0214]: with Ranking-based Active Learning, the system is programmed to re-rank data; rather than predicting which record should be labeled in the context of a specific loop, the system is programmed to continuously re-rank records in the dataset based on qualitative aspects of the records (e.g., how effective the dataset is in training the model), which allows the system to select data not only for loop n, but the subsequent ones most effective in training the model as well) (Martin, ¶ [0029]: after a step in the path, the labeling platform will have more information about the actual confidence and costs so far in the execution path and a redetermination of candidate paths can be performed help optimize the path from the current point in the execution path forward, whether or not the expectations for the current point in the path have been met; ¶ [0041]: the CDW dynamically determines a cost-optimized execution path to meet a target confidence threshold for a labeling request; ¶ [0118]: CDW can be configured dynamically determine an execution path to optimize the confidence of the overall labeling result that incorporates the individual labeling results obtained so far for the labeling request; ¶ [0186]: if there are cost constraints (temporal, monetary or other cost) and/or confidence constraints, workflow orchestrator 710 performs a path search to find the optimal set of labelers (step 1114); ¶¶ [0197]-[0204]: workflow orchestrator 710 can pick the optimal path from the possibly multiple solutions found that meet the constraints; workflow orchestrator can optimize the path according to one or more of: 1) Cost optimized; 2) Time optimized; 3) Confidence optimized; combinations of these can be used by optimizing a set of confidence constraints for a range, then optimizing the subset for a different range on a different variable) (Li, Abstract in Page 1: in the deployment of deep neural models, how to effectively and automatically find feasible deep models under diverse design objectives is fundamental. Most existing neural architecture search (NAS) methods utilize surrogates to predict the detailed performance (e.g., accuracy and model size) of a candidate architecture during the search, which however is complicated and inefficient; in contrast, we aim to learn an efficient Pareto classifier to simplify the search process of NAS by transforming the complex multi-objective NAS task into a simple Pareto-dominance classification task; propose a classification wise Pareto evolution approach for one-shot NAS, where an online classifier is trained to predict the dominance relationship between the candidate and constructed reference architectures, instead of using surrogates to fit the objective functions; the main contribution of this study is to change supernet adaption into a Pareto classifier; besides, we design two adaptive schemes to select the reference set of architectures for constructing classification boundary and regulate the rate of positive samples over negative ones, respectively; Section I of Pages 1-2: NEURAL architecture search (NAS) is shown to be promising for the automatic design of task-specific deep neural networks (DNNs), instead of the traditional manual design based on extensive human expertise [1], [2], [3], [4]; consequently, NAS has received a surge of attention from the community of deep learning, largely owing to its superiority in optimizing the architecture and weights of DNN [2]; in other words, NAS can obtain desirable architectures as human experts do, or even far more innovative architectures [3]; although EA-based NAS is gradient free and insensitive to the complexity of the objectives [3], it suffers complicated and costly evaluation of candidate architectures on specific tasks; another issue of NAS is how to effectively optimize multiple conflicting design objectives (e.g., accuracy and model size) in solving practical applications; although many multi-objective selection strategies (e.g., the non-dominate sorting strategy [18]) have been integrated into NAS to drive the search towards the Pareto front over different objectives [19], there is still a big gap to improve the effectiveness [20]; specifically, if the objectives are not correctly evaluated, the Pareto dominance relationships between architectures will be easily misjudged, and the subsequent search process of NAS will be misled [21]; specifically, the classifier aims to predict the dominance relationship of each candidate architecture over a set of reference points (or architectures), and classify the candidate architectures into two (good and poor) classes; the samples in the good class are the points which either dominate or are non-dominant with the reference points; the poor class contains the samples dominated by the reference points; in this way, the complex architecture evaluation task is transformed into a Pareto-dominance classification task; besides, we design an adaptive clustering-based selection method and an α-dominance evaluation strategy to produce effective samples to avoid the class imbalance problem that the number of samples in one class is far less or more than that of another class; hence, high accuracy of the classification will be achieved; Section 3 with FIG. 1 and Algorithms 1-7 in Pages 3-8: Fig. 1 shows the framework of CENAS, which includes four main components: a multi-objective search routine, an ensemble Pareto classifier, a set of reference points, and an adaptive supernet; especially, the Pareto classifier is used to divide offspring architectures into two (good and poor) classes, while the set of reference points are used to construct training samples for classifier; an initial population of N architectures (subnets) is generated via the uniform sampling from the search space; then, the performance objectives (f1,…,fm) of each architecture in P are evaluated using weights inherited from the supernet on the validation set; a set of reference points SR is required to be In this process, an angle-based cluster strategy is used to make the reference points adapt to the shape of the Pareto front; moreover, a relaxed dominance strategy is employed to adjust the rate of positive and negative samples, so as to avoid the class imbalance problem selected from Arc to construct the classification boundary; the architectures in the archive are categorized into two classes (class1-good/class2-poor) according to the classification boundary (see Algorithm 3); then these samples are divided into the training data set and test data set for the training of a Pareto classifier, which is based on the ensemble of support vector machines (SVMs) (see Algorithm 4 and Algorithm 5); reproduction operators like crossover and mutation are performed on the parent solutions to create offspring solutions; promising architectures from the offspring are selected by the proposed classifier as the parent individuals of the next generation (see Algorithm 6); the archive Arc aims to keep track of promising architectures gained from the evolutionary search process so far, and it is updated using the obtained population P; then, the task-specific supernet is trained using the subnets sampled from architectures in the archive (see Algorithm 7); from the archive, we can obtain desired DNNs having the best trade-off between different objectives; present a densely-connected-like search space as shown in Fig. 2 (top left); specifically, a series of cells, i.e., the building blocks of an architecture, are connected in a feed-forward manner; two types of cells are designed: the normal cells aim to preserve the input size of data (or feature map), and the reduction cells to reduce the input spatial size; our goal is to find a suitable routing-like path between these cells to form a final architecture; in order to train a classifier to distinguish the candidate architectures, we select a set of reference points (architectures) as classification boundary to divide the samples into two (i.e., good and poor) classes; Apparently, the selected reference points determine the labeling of samples; Fig. 3 gives an example to show the impact of the position of reference points on the classification; design an adaptive clustering-based approach to select reference points from the obtained architectures. The main idea is to use clustering to help detect the distribution of the obtained solutions and guide the selection of reference points in line with the Pareto front; the details are given in Algorithm 2; non-domination sorting (Line 1): the nondominated fronts (PF1, PF2, …) of solutions are obtained by conducting conventional non-dominated sorting on the population from scratch [18], where the solutions in each front are nondominated with each other. The nondominated sorting is detailed in Section S1 of the supplementary material; to alleviate the computation cost of evaluating subnets under multiple objectives, we build an online Pareto classifier, which is to directly predict the Pareto-dominance classification of a subnet without training; Pareto dominance comparison is used to construct the labels of samples in the following way: if a solution is dominated by the set of reference points, then it will be labeled as a negative sample; otherwise it is a positive one; in the labeling process, the use of original Pareto-dominance may result in a class imbalance phenomenon; to handle this issue, we use α-domination criterion [37] to re-label samples (Lines 13-28 in Algorithm 3); the α-domination criterion. The α-domination is a relaxed Pareto-dominance by introducing a parameter α to control the dominated region; In Fig. 4, it can be seen that the α-domination is able to regulate the dominated region of solutions, and control the number of dominated solutions).
Claim 18
Prendki in view of Martin discloses all the elements as stated in Claim 16 and further discloses wherein determining the first automatic labeling procedure based on the Pareto dominance relation comprises: computing one or more second numbers corresponding to the set of automatic labeling tasks by at least: for each automatic labeling procedure in the set of automatic labeling procedures, computing a second number of rest automatic labeling procedures dominated by each automatic labeling procedure in the set of the automatic labeling procedures based on the Pareto dominance relation, where the rest automatic labeling procedures include remaining automatic labeling procedures after selecting any one of the automatic labeling procedure from the set of automatic labeling procedures; sorting a priority of the set of automatic labeling procedures based on the one or more second numbers; determining the first automatic labeling procedure as an automatic labeling procedure with the highest second number in the one or more second numbers ((Prendki, ¶ [0050]: efficiently compare the capabilities, availability, performance, and costs of various labeling providers; ¶ [0057]:recommendation computer system 306 is programmed using the real-time evaluation instructions 308 to query the provider metrics database 310 and to select, automatically and based upon programmed recommendation or evaluation algorithms, one or more of the label service providers 314, 316, 318 that is optimal to perform one or more labeling tasks on the user dataset; ¶¶ [0114]-[0154]: collect various information of the labeling partners, which may be used by the optimization system to find the most suitable labeling partner for a customer-requested task; examples of such information that the system collects of the labeling partners may include the following: list of the types of tasks that the partner can deliver on (depends on tooling and labor workforce); available clearance; annotator information (e.g., type of tasks annotators are capable of working on); historical performance of the annotators (e.g., speed and accuracy); desired workload of annotators; and real-time availability of annotators; compute and maintain various metrics of the labeling providers; examples of such metrics may include a) average labeling quality (e.g., accuracy) for various types of tasks; b) average labeling speed for various types of tasks; c) typical availability; d) average price, and deviation from standard price; e) time since last offered task; f) time of idle (e.g., no assigned task); g) time taken to resolve labeling issues identified by the quality control process; and number of iterations required to solve a labeling issue; h) response time to the system's inquiry/notifications; and i) acceptance/rejection/ignore statistics of labeling tasks; utilize a multi-objective optimization process to find a labeling partner that is most likely to produce the highest quality of work; examples of the factors considered during the optimization process include, but are not limited to, the following: a) turn-around time (e.g., on average how long does the labeling partner need to complete and return a task); b) labeling quality (e.g., how accurate is the labeled data); c) average cost (e.g., the average cost charged by the labeling partner for various types of tasks); d) revenue (e.g., how much did the labeling partner earn for a given period of time); e) number of, and type of, tasks previously offered to the labeling partner; f) number of labeling partners within the computer-implemented labeling marketplace; and g) customer's preferences/criteria; the optimization process may be programmed using any of: the Dantzig simplex algorithm, extensions or variants thereof; combinatorial algorithms; or quantum optimization algorithms; Claim 10: the identifying by multi-objective optimization based on: an average time required by the labeling partners to complete and return a task; an average labeling accuracy of the labeling partners; an average price of the labeling partners; an average revenue of the labeling partners over a given period of time; a number of tasks previously offered to the labeling partners; and types of tasks previously offered to the labeling partners; ¶¶ [0174] and [0228]-[0229]: the system may end the experiment if it determines that the model has reached its optimal performance; the system determines that the model has reached its optimal performance if the portions of the data identified by the data curation engine as being effective at training the model have all been labeled and only the portions that are identified as redundant or irrelevant in training the model are left; ¶ [0206]: provides universal labeling instructions that are optimized for all labeling partners within the computer-implemented labeling marketplace; the universal labeling instructions are designed to be universally compatible with all of the labeling partners within the computer-implemented labeling marketplace by considering each of the labeling partners' capabilities at both the provider-level and annotator-level and the tooling available to the partners; ¶ [0220]: applies an optimization process that evaluates real-time and historical metrics of all the labeling partners remaining in the selection pool; ¶ [0214]: with Ranking-based Active Learning, the system is programmed to re-rank data; rather than predicting which record should be labeled in the context of a specific loop, the system is programmed to continuously re-rank records in the dataset based on qualitative aspects of the records (e.g., how effective the dataset is in training the model), which allows the system to select data not only for loop n, but the subsequent ones most effective in training the model as well) (Martin, ¶ [0029]: after a step in the path, the labeling platform will have more information about the actual confidence and costs so far in the execution path and a redetermination of candidate paths can be performed help optimize the path from the current point in the execution path forward, whether or not the expectations for the current point in the path have been met; ¶ [0041]: the CDW dynamically determines a cost-optimized execution path to meet a target confidence threshold for a labeling request; ¶ [0118]: CDW can be configured dynamically determine an execution path to optimize the confidence of the overall labeling result that incorporates the individual labeling results obtained so far for the labeling request; ¶ [0186]: if there are cost constraints (temporal, monetary or other cost) and/or confidence constraints, workflow orchestrator 710 performs a path search to find the optimal set of labelers (step 1114); ¶¶ [0197]-[0204]: workflow orchestrator 710 can pick the optimal path from the possibly multiple solutions found that meet the constraints; workflow orchestrator can optimize the path according to one or more of: 1) Cost optimized; 2) Time optimized; 3) Confidence optimized; combinations of these can be used by optimizing a set of confidence constraints for a range, then optimizing the subset for a different range on a different variable) (Li, Abstract in Page 1: in the deployment of deep neural models, how to effectively and automatically find feasible deep models under diverse design objectives is fundamental. Most existing neural architecture search (NAS) methods utilize surrogates to predict the detailed performance (e.g., accuracy and model size) of a candidate architecture during the search, which however is complicated and inefficient; in contrast, we aim to learn an efficient Pareto classifier to simplify the search process of NAS by transforming the complex multi-objective NAS task into a simple Pareto-dominance classification task; propose a classification wise Pareto evolution approach for one-shot NAS, where an online classifier is trained to predict the dominance relationship between the candidate and constructed reference architectures, instead of using surrogates to fit the objective functions; the main contribution of this study is to change supernet adaption into a Pareto classifier; besides, we design two adaptive schemes to select the reference set of architectures for constructing classification boundary and regulate the rate of positive samples over negative ones, respectively; Section I of Pages 1-2: NEURAL architecture search (NAS) is shown to be promising for the automatic design of task-specific deep neural networks (DNNs), instead of the traditional manual design based on extensive human expertise [1], [2], [3], [4]; consequently, NAS has received a surge of attention from the community of deep learning, largely owing to its superiority in optimizing the architecture and weights of DNN [2]; in other words, NAS can obtain desirable architectures as human experts do, or even far more innovative architectures [3]; although EA-based NAS is gradient free and insensitive to the complexity of the objectives [3], it suffers complicated and costly evaluation of candidate architectures on specific tasks; another issue of NAS is how to effectively optimize multiple conflicting design objectives (e.g., accuracy and model size) in solving practical applications; although many multi-objective selection strategies (e.g., the non-dominate sorting strategy [18]) have been integrated into NAS to drive the search towards the Pareto front over different objectives [19], there is still a big gap to improve the effectiveness [20]; specifically, if the objectives are not correctly evaluated, the Pareto dominance relationships between architectures will be easily misjudged, and the subsequent search process of NAS will be misled [21]; specifically, the classifier aims to predict the dominance relationship of each candidate architecture over a set of reference points (or architectures), and classify the candidate architectures into two (good and poor) classes; the samples in the good class are the points which either dominate or are non-dominant with the reference points; the poor class contains the samples dominated by the reference points; in this way, the complex architecture evaluation task is transformed into a Pareto-dominance classification task; besides, we design an adaptive clustering-based selection method and an α-dominance evaluation strategy to produce effective samples to avoid the class imbalance problem that the number of samples in one class is far less or more than that of another class; hence, high accuracy of the classification will be achieved; Section 3 with FIG. 1 and Algorithms 1-7 in Pages 3-8: Fig. 1 shows the framework of CENAS, which includes four main components: a multi-objective search routine, an ensemble Pareto classifier, a set of reference points, and an adaptive supernet; especially, the Pareto classifier is used to divide offspring architectures into two (good and poor) classes, while the set of reference points are used to construct training samples for classifier; an initial population of N architectures (subnets) is generated via the uniform sampling from the search space; then, the performance objectives (f1,…,fm) of each architecture in P are evaluated using weights inherited from the supernet on the validation set; a set of reference points SR is required to be In this process, an angle-based cluster strategy is used to make the reference points adapt to the shape of the Pareto front; moreover, a relaxed dominance strategy is employed to adjust the rate of positive and negative samples, so as to avoid the class imbalance problem selected from Arc to construct the classification boundary; the architectures in the archive are categorized into two classes (class1-good/class2-poor) according to the classification boundary (see Algorithm 3); then these samples are divided into the training data set and test data set for the training of a Pareto classifier, which is based on the ensemble of support vector machines (SVMs) (see Algorithm 4 and Algorithm 5); reproduction operators like crossover and mutation are performed on the parent solutions to create offspring solutions; promising architectures from the offspring are selected by the proposed classifier as the parent individuals of the next generation (see Algorithm 6); the archive Arc aims to keep track of promising architectures gained from the evolutionary search process so far, and it is updated using the obtained population P; then, the task-specific supernet is trained using the subnets sampled from architectures in the archive (see Algorithm 7); from the archive, we can obtain desired DNNs having the best trade-off between different objectives; present a densely-connected-like search space as shown in Fig. 2 (top left); specifically, a series of cells, i.e., the building blocks of an architecture, are connected in a feed-forward manner; two types of cells are designed: the normal cells aim to preserve the input size of data (or feature map), and the reduction cells to reduce the input spatial size; our goal is to find a suitable routing-like path between these cells to form a final architecture; in order to train a classifier to distinguish the candidate architectures, we select a set of reference points (architectures) as classification boundary to divide the samples into two (i.e., good and poor) classes; Apparently, the selected reference points determine the labeling of samples; Fig. 3 gives an example to show the impact of the position of reference points on the classification; design an adaptive clustering-based approach to select reference points from the obtained architectures. The main idea is to use clustering to help detect the distribution of the obtained solutions and guide the selection of reference points in line with the Pareto front; the details are given in Algorithm 2; non-domination sorting (Line 1): the nondominated fronts (PF1, PF2, …) of solutions are obtained by conducting conventional non-dominated sorting on the population from scratch [18], where the solutions in each front are nondominated with each other. The nondominated sorting is detailed in Section S1 of the supplementary material; to alleviate the computation cost of evaluating subnets under multiple objectives, we build an online Pareto classifier, which is to directly predict the Pareto-dominance classification of a subnet without training; Pareto dominance comparison is used to construct the labels of samples in the following way: if a solution is dominated by the set of reference points, then it will be labeled as a negative sample; otherwise it is a positive one; in the labeling process, the use of original Pareto-dominance may result in a class imbalance phenomenon; to handle this issue, we use α-domination criterion [37] to re-label samples (Lines 13-28 in Algorithm 3); the α-domination criterion. The α-domination is a relaxed Pareto-dominance by introducing a parameter α to control the dominated region; In Fig. 4, it can be seen that the α-domination is able to regulate the dominated region of solutions, and control the number of dominated solutions)).
Allowable Subject Matter
Claims 13 and 20 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 101 and U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
Claim 13
Prendki in view of Martin discloses all the elements as stated in Claims 1-12 and 14-19; ZOU discloses all the elements as stated in Claim 10; and Li discloses all the elements as stated in Claims 16-18.
McKay et al. (US 2021/0192394 A1, pub. date: 06/24/2021) discloses in Abstract that systems, methods and products for optimization of a machine learning labeler; labeling requests are received and corresponding label inferences are generated using a champion model; a portion of the labeling requests and corresponding inferences is selected for use as training data, and labels are generated for the selected requests, thereby producing corresponding augmented results; a first portion of the augmented results are provided as training data to an experiment coordinator, which then trains one or more challenger models using these augmented results; a second portion of the augmented results is provided as evaluation data to a model evaluator, which evaluates the performance of the challenger models and the champion model; and if one of the challenger models has higher performance than the champion model, the model evaluator promotes the challenger model to replace the champion model. McKay further discloses in ¶¶ [0122]-[0123] with FIG. 10 that an ML labeler that includes multiple internal ML labelers; a splitter 1004 is implemented in the conditioning layer on the input pipe and training pipe; here the splitter splits a request to label an image (or image training data) into requests to constituent ML labelers 1002a, 1002b, 1002c, 1002d, where each constituent ML labeler is trained for a particular product category; e.g., splitter 1004 routes the labeling request to i) labeler 1002a to label the image with any tools that labeler 1002a detects in the image, ii) labeler 1002b to label the image with any vehicles that labeler 1002b detects in the image, iii) labeler 1002c to label the image with any clothing items that labeler 1002c detects in the image, and iv) labeler 1002d to label the image with any food items that labeler 1002d detects in the image; a label and confidence aggregator 1006 is implemented in the conditioning layer on the output pipe to aggregate the inferences and confidences output by labelers 1002a, 1002b, 1002c, and 1002d to determine the label(s) and confidence(s) applicable to the image; thus, a conditioning component may result in fan-in and fan-out conditions in a directed graph; FIG. 10 involves two fan-out points and one fan-in point: a) labeling request fan-out to route the same labeling request to each constituent product area labeler; b) labeling result fan-in to assemble labeling results from each constituent labeler into an overall labeling result; and c) training data fan-out to split the training data labels by product type, and route the appropriate label sets to the correct constituent labelers; splitting or slicing can be achieved by label splitter components implemented in the respective conditioning pipelines; fan-out can be configured by linking several labelers' request pipes to a single result pipe of a conditioning component; fan-in can be achieved by linking multiple output pipes to a single input pipe of an aggregator conditioning component; and the aggregator can be configured with an aggregation key identifier identifying which constituent data should be aggregated, a template specifying how to combine the inferences from multiple labelers and an algorithm for aggregating the confidences.
DALLI et al. (US 2022/0398460 A1, pub. date: 12/15/2022) discloses in Abstract that an exemplary model search may provide optimal explainable models based on a dataset; identify features from a training dataset, and may map feature costs to the identified features; the search space may be sampled to generate initial or seed candidates, which may be chosen based on one or more objectives and/or constraints; the candidates may be iteratively optimized until an exit condition is met; the optimization may be performed by an external optimizer; the external optimizer may iteratively apply constraints to the candidates to quantify a fitness level of each of the seed candidates; the fitness level may be based on the constraints and objectives. The candidates may be a set of data, or may be trained to form explainable models; the external optimizer may optimize the explainable models until the exit conditions are met. DALLI further disclose in ¶¶ [0010], [0262], [0333], [0355]-[0358], [0374], that Multi-objective optimization (MOO) may be used to solve problems that have competing objectives; some MOO methods weight the different measures according to their importance, essentially scalarizing the problem and converting the original problem into a single-objective optimization problem; another technique-the E-constraint method-minimizes one objective while keeping the other objectives at user-set values or subject to inequality constraints; and other techniques may also be employed for these purposes, including the weighted metric method and evolutionary algorithm MOO techniques; MOO techniques may involve an iterative improvement of what may be referred to as the Pareto front, the set of Pareto-optimal candidates for which one cannot find better alternatives without resorting to a worse tradeoff between search parameters; FIG. 1 illustrates the Pareto front plot for an exemplary multi-objective optimization scenario with two variables: performance (or accuracy) 601 and explainability 602; successive generations of candidates may improve the Pareto front 603-pushing it farther to the right and upwards; as such, the Pareto front may include any candidates which are optimal; there are various philosophies for the solution of MOO problems; in the simplest of cases, MOO may be used as a selection tool to help pick the best candidate from among several candidates, leaving the generation of candidates for another, separate algorithm; in other processes, MOO may be more tightly integrated with the candidate-generating algorithm; Aside from PKI, an exemplary embodiment may also exploit learned representations during the search process so that new candidate architectures are informed by the latest Pareto-optimal candidates, even if the architectures do not perfectly match; this is a type of transfer learning and may be achieved by morphing existing architectures and reusing previous parameters during the candidate generation; new candidates may be regenerated 1212 via, for example, recombination (if using GA optimization) or other evolutionary techniques from Pareto-optimal candidates; each iteration of the external optimizer may update the Pareto front 1210; the Pareto front may refer to the set of candidates that are not Pareto dominated; including the set of candidates whose fitness cannot be improved without losing the tradeoff, i.e., without degrading the scalar fitness of some other objective; thus, the Pareto front may be updated by identifying which candidates optimize the available metrics based on the values closest to the objectives and which are bound by the constraints; the Pareto front may be useful in MOO, because it may embody the "best" current selection of candidates; future candidates may be regenerated from parent candidates among the Pareto front; candidates on the Pareto front are often referred to as potential candidates due to the MOO process.
However, closest arts of records, as discussed above, singly or in combination do not teach or suggest at least following features "wherein obtaining the set of automatic labeling procedures based on the task type of the target labeling task comprises: selecting, based at least in part on a predefined feature extraction model criteria and a classifier model criteria, X feature extraction models and Y classifier models respectively based on the task type of the target labeling task, X and Y being positive integers; obtaining a set of automatic labeling procedures from permutating and combining any one feature extraction model of the X feature extraction models with any one classifier model of the Y classifier models" when combining all other limitations of the claim as a whole.
Claim 20
Prendki in view of Martin discloses all the elements as stated in Claims 1-12 and 14-19; ZOU discloses all the elements as stated in Claim 10; Li discloses all the elements as stated in Claims 16-18; McKay discloses all the elements as stated in 13; and DALLI discloses all the elements as stated in 13.
However, closest arts of records, as discussed above, singly or in combination do not teach or suggest at least following features" wherein determining that the first automatic labeling procedure fails to satisfy the first quality evaluation index further comprises: assigning the automatic labeling data of the target labeling task as unlabeled data; updating the first automatic labeling procedure until a Pareto target value of the first automatic labeling procedure converges, wherein the unlabeled data are performed with following extraction and labeling operations comprising: performing random extraction on the unlabeled data to obtain randomly extracted data; computing a contribution value of respective data items in the randomly extracted data to a multi-objective optimization model, the multi-objective optimization model being provided for computing a first Pareto optimal target value of the first automatic labeling procedure; selecting data items from the randomly extracted data based on contribution values of the respective data items in the randomly extracted data, to form extraction data with high contribution value; manually labeling the extraction data with a high contribution value to obtain a corresponding manual labeling result; computing a second Pareto optimal target value of the first automatic labeling procedure based on the extraction data with the high contribution value and the corresponding manual labeling result; updating the first automatic labeling procedure based on the extraction data with the high contribution value and the corresponding manual labeling result; and updating the unlabeled data based on the extraction data with the high contribution value and the corresponding manual labeling result; and continuing to perform the extraction and labeling operations based on the unlabeled data updated until the second Pareto target value converges" when combining all other limitations of the claim as a whole.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HWEI-MIN LU whose telephone number is (313)446-4913. The examiner can normally be reached Mon - Fri: 9:00 AM - 6:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela D. Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HWEI-MIN LU/Primary Examiner, Art Unit 2142