DETAILED ACTION
This communication is in response to Application No. 18/636,730 filed on April 16, 2024
in which Claims 1-20 are presented for examination.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 12 is objected to because of the following informality: claim 12 recites "The medium of claim 114," which appears to be a typographical error, as there is no claim 114 in the application. It is suggested that claim 12 be amended to depend on claim 11. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 1-20 are rejected under 35 U.S.C. 101 because these claimed inventions are directed to an
abstract idea without significantly more.
Regarding Claim 1:
Step 1: Claim 1 is a method type claim. Therefore, Claims 1-7 fall within one of the four statutory
categories (i.e., process, machine, manufacture, or composition of matter).
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance
of the limitation in the mind but for the recitation of generic computer components, then it falls within
the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable
interpretation, covers performance of the limitation by mathematical calculation but for the recitation
of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract
ideas.
updating a fidelity metric associated with each of the one or more human evaluators based on the evaluation (mental process – updating a fidelity metric based on an evaluation may be performed mentally by a user observing/analyzing the received evaluation and accordingly using judgment/evaluation to adjust a value associated with each evaluator based on said analysis)
determining a cumulative ranking of the answer with respect to the question according to the evaluation and the updated fidelity metric of each of the one or more human evaluators (mental process - determining a cumulative ranking may be performed mentally by a user reading/analyzing the evaluations and the fidelity metrics and accordingly using judgment/evaluation to arrive at a ranking based on said analysis)
updating a fidelity attribute associated with the machine expert based on the cumulative ranking, wherein the fidelity attribute is indicative of an ability of the machine expert in answering questions in the subject matter (mental process - updating a fidelity attribute based on the cumulative ranking may be performed mentally by a user analyzing the cumulative ranking and accordingly using judgment/evaluation to adjust a value indicative of the machine expert's ability based on said analysis)
generating feedback based on the answer, the question, the cumulative ranking of the answer with respect to the question, and the updated fidelity attribute of the machine expert (mental process - generating feedback may be performed mentally or using pen and paper by a user reviewing/analyzing the answer, question, ranking, and fidelity attribute and accordingly using judgment/evaluation to formulate/write feedback based on said analysis)
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
receiving, from one or more human evaluators, evaluation directed to an answer […] (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g))
[…] automatically generated by a machine expert in a question & answer (Q&A) system in response to a question related to a subject matter based on a reference from a source (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine expert (generative AI) to generate answers without significantly more)
sending the feedback to the Q&A system for adapting the Q&A system (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g))
Step 2B: The claim does not include additional elements considered individually and in combination that
are sufficient to amount to significantly more than the judicial exception.
receiving, from one or more human evaluators, evaluation directed to an answer […] (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer)
[…] automatically generated by a machine expert in a question & answer (Q&A) system in response to a question related to a subject matter based on a reference from a source (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine expert (generative AI) to generate answers without significantly more)
sending the feedback to the Q&A system for adapting the Q&A system (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer)
For the reasons above, Claim 1 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 2 - 7. The additional limitations of the dependent claims are addressed below.
Regarding Claim 2:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 2 depends on.
Step 2A Prong 2 & Step 2B:
wherein the Q&A system includes a plurality of machine experts for automatically generating answers to questions […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine expert (generative AI) to generate answers without significantly more)
[…] wherein, for each question asked, at least some of the plurality of machine experts are selected for providing an answer to the question and the selection is based, at least partially, on the fidelity attribute associated with each of the plurality of machine experts space (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the plurality of machine experts are selected for providing an answer to the question and the selection is based, at least partially, on the fidelity attribute associated with each of the plurality of machine experts space does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 3:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 3 depends on.
identifying a ranking for the answer provided by the human evaluator from the evaluation (mental process - identifying a ranking for the answer from the evaluation may be performed mentally by a user reading/analyzing the evaluation and accordingly using judgment/evaluation to identify the ranking based on said analysis)
determining a number of other rankings from remainder of the one or more human evaluators that are consistent with the ranking (mental process - determining a number of other rankings that are consistent with the ranking may be performed mentally or using pen and paper by a user reviewing/analyzing the rankings from the other evaluators and accordingly using judgment/evaluation to count those consistent with the ranking based on said analysis)
determining a parameter based on the number of rankings from others (mental process - determining a parameter based on the number of rankings from others may be performed mentally by a user analyzing the number of consistent rankings and accordingly using judgment/evaluation to determine a parameter based on said analysis)
updating an existing fidelity metric associated with the human evaluator based on the parameter to generate the updated fidelity metric for the human evaluator (mental process - updating an existing fidelity metric based on the parameter may be performed mentally or using pen and paper by a user analyzing the parameter and accordingly using judgment/evaluation to adjust the fidelity metric associated with the evaluator based on said analysis)
Step 2A Prong 2 & Step 2B:
Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 4:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 4 depends on.
identifying a ranking from the evaluation from each of the one or more human evaluators (mental process - identifying a ranking from the evaluation from each evaluator may be performed mentally by a user reading/analyzing each evaluation and accordingly using judgment/evaluation to identify the ranking based on said analysis)
weighing the ranking of each of the one or more human evaluators based on the updated fidelity metric thereof to generate a weighted ranking for the human evaluator (mental process - weighing the ranking of each evaluator based on the updated fidelity metric may be performed mentally by a user analyzing the ranking and the fidelity metric and accordingly using judgment/evaluation to generate a weighted ranking based on said analysis)
obtaining an integrated ranking for the answer based on the weighted ranking of each of the one or more human evaluators (mental process - obtaining an integrated ranking based on the weighted rankings may be performed mentally by a user analyzing the weighted rankings and accordingly using judgment/evaluation to obtain an integrated ranking based on said analysis)
determining the cumulative ranking of the answer based on the existing ranking and the integrated ranking for the answer (mental process - determining the cumulative ranking based on the existing ranking and the integrated ranking may be performed mentally by a user analyzing the existing and integrated rankings and accordingly using judgment/evaluation to determine the cumulative ranking based on said analysis)
Step 2A Prong 2 & Step 2B:
retrieving an existing ranking for the answer for the question (adding insignificant extra-solution activity to the judicial exception, mere data retrieval/gathering of the existing ranking that is then analyzed – see MPEP 2106.05(g); and, retrieving data is a well-understood, routine, and conventional function when claimed in a merely generic manner, as it is here, supporting a conclusion under Berkheimer – see MPEP 2106.05(d))
accessing the updated fidelity metric for each of the one or more human evaluators (adding insignificant extra-solution activity to the judicial exception, mere data retrieval/gathering of the fidelity metric that is then used in the analysis – see MPEP 2106.05(g); and, accessing data is a well-understood, routine, and conventional function when claimed in a merely generic manner, as it is here, supporting a conclusion under Berkheimer – see MPEP 2106.05(d))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 5:
Step 2A Prong 1: See the rejection of Claim 4 above, which Claim 5 depends on.
modifying the existing fidelity attribute based on the cumulative ranking of the answer (mental process - modifying the existing fidelity attribute based on the cumulative ranking may be performed mentally by a user analyzing the cumulative ranking and accordingly using judgment/evaluation to adjust the fidelity attribute based on said analysis)
generating an updated fidelity attribute based on the modified fidelity attribute for the machine expert (mental process - generating an updated fidelity attribute based on the modified fidelity attribute may be performed mentally or using pen and paper by a user analyzing the modified fidelity attribute and accordingly using judgment/evaluation to generate the updated fidelity attribute based on said analysis)
Step 2A Prong 2 & Step 2B:
retrieving an existing fidelity attribute associated with the machine expert (adding insignificant extra-solution activity to the judicial exception, mere data retrieval/gathering of the existing ranking that is then analyzed – see MPEP 2106.05(g); and, retrieving data is a well-understood, routine, and conventional function when claimed in a merely generic manner, as it is here, supporting a conclusion under Berkheimer – see MPEP 2106.05(d))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 4. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 6:
Step 2A Prong 1: See the rejection of Claim 1 above, which Claim 6 depends on.
Step 2A Prong 2 & Step 2B:
wherein the feedback further includes at least one of an alternative answer in place of the answer provided by one of the one or more human evaluators, an alternative reference that supports the alternative answer, and an alternative source to access the alternative reference (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the feedback further includes at least one of an alternative answer in place of the answer provided by one of the one or more human evaluators, an alternative reference that supports the alternative answer, and an alternative source to access the alternative reference does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 1. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 7:
Step 2A Prong 1: See the rejection of Claim 6 above, which Claim 7 depends on.
extracting, from the feedback, information related to the alternative reference and the alternative source (mental process - extracting information related to the alternative reference and the alternative source from the feedback may be performed mentally by a user reading/analyzing the feedback and accordingly using judgment/evaluation to identify and extract the relevant information based on said analysis)
modifying an archive storing references from different sources based on the alternative reference and the alternative source (mental process - modifying an archive storing references based on the alternative reference and the alternative source may be performed mentally or using pen and paper by a user analyzing the alternative reference and source and accordingly using judgment/evaluation to update the stored references based on said analysis)
Step 2A Prong 2 & Step 2B:
Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 8:
Step 1: Claim 8 is a non-transitory medium type claim. Therefore, Claims 8 - 14 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter).
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance
of the limitation in the mind but for the recitation of generic computer components, then it falls within
the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable
interpretation, covers performance of the limitation by mathematical calculation but for the recitation
of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract
ideas.
updating a fidelity metric associated with each of the one or more human evaluators based on the evaluation (mental process – updating a fidelity metric based on an evaluation may be performed mentally by a user observing/analyzing the received evaluation and accordingly using judgment/evaluation to adjust a value associated with each evaluator based on said analysis)
determining a cumulative ranking of the answer with respect to the question according to the evaluation and the updated fidelity metric of each of the one or more human evaluators (mental process - determining a cumulative ranking may be performed mentally by a user reading/analyzing the evaluations and the fidelity metrics and accordingly using judgment/evaluation to arrive at a ranking based on said analysis)
updating a fidelity attribute associated with the machine expert based on the cumulative ranking, wherein the fidelity attribute is indicative of an ability of the machine expert in answering questions in the subject matter (mental process - updating a fidelity attribute based on the cumulative ranking may be performed mentally by a user analyzing the cumulative ranking and accordingly using judgment/evaluation to adjust a value indicative of the machine expert's ability based on said analysis)
generating feedback based on the answer, the question, the cumulative ranking of the answer with respect to the question, and the updated fidelity attribute of the machine expert (mental process - generating feedback may be performed mentally or using pen and paper by a user reviewing/analyzing the answer, question, ranking, and fidelity attribute and accordingly using judgment/evaluation to formulate/write feedback based on said analysis)
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
receiving, from one or more human evaluators, evaluation directed to an answer […] (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g))
[…] automatically generated by a machine expert in a question & answer (Q&A) system in response to a question related to a subject matter based on a reference from a source (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine expert (generative AI) to generate answers without significantly more)
sending the feedback to the Q&A system for adapting the Q&A system (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g))
Step 2B: The claim does not include additional elements considered individually and in combination that
are sufficient to amount to significantly more than the judicial exception.
receiving, from one or more human evaluators, evaluation directed to an answer […] (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer)
[…] automatically generated by a machine expert in a question & answer (Q&A) system in response to a question related to a subject matter based on a reference from a source (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine expert (generative AI) to generate answers without significantly more)
sending the feedback to the Q&A system for adapting the Q&A system (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer)
For the reasons above, Claim 8 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 9 - 14. The additional limitations of the dependent claims are addressed below.
Regarding Claim 9:
Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 9 depends on.
Step 2A Prong 2 & Step 2B:
wherein the Q&A system includes a plurality of machine experts for automatically generating answers to questions […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine expert (generative AI) to generate answers without significantly more)
[…] wherein, for each question asked, at least some of the plurality of machine experts are selected for providing an answer to the question and the selection is based, at least partially, on the fidelity attribute associated with each of the plurality of machine experts space (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the plurality of machine experts are selected for providing an answer to the question and the selection is based, at least partially, on the fidelity attribute associated with each of the plurality of machine experts space does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 8. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 10:
Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 10 depends on.
identifying a ranking for the answer provided by the human evaluator from the evaluation (mental process - identifying a ranking for the answer from the evaluation may be performed mentally by a user reading/analyzing the evaluation and accordingly using judgment/evaluation to identify the ranking based on said analysis)
determining a number of other rankings from remainder of the one or more human evaluators that are consistent with the ranking (mental process - determining a number of other rankings that are consistent with the ranking may be performed mentally or using pen and paper by a user reviewing/analyzing the rankings from the other evaluators and accordingly using judgment/evaluation to count those consistent with the ranking based on said analysis)
determining a parameter based on the number of rankings from others (mental process - determining a parameter based on the number of rankings from others may be performed mentally by a user analyzing the number of consistent rankings and accordingly using judgment/evaluation to determine a parameter based on said analysis)
updating an existing fidelity metric associated with the human evaluator based on the parameter to generate the updated fidelity metric for the human evaluator (mental process - updating an existing fidelity metric based on the parameter may be performed mentally or using pen and paper by a user analyzing the parameter and accordingly using judgment/evaluation to adjust the fidelity metric associated with the evaluator based on said analysis)
Step 2A Prong 2 & Step 2B:
Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 11:
Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 11 depends on.
identifying a ranking from the evaluation from each of the one or more human evaluators (mental process - identifying a ranking from the evaluation from each evaluator may be performed mentally by a user reading/analyzing each evaluation and accordingly using judgment/evaluation to identify the ranking based on said analysis)
weighing the ranking of each of the one or more human evaluators based on the updated fidelity metric thereof to generate a weighted ranking for the human evaluator (mental process - weighing the ranking of each evaluator based on the updated fidelity metric may be performed mentally by a user analyzing the ranking and the fidelity metric and accordingly using judgment/evaluation to generate a weighted ranking based on said analysis)
obtaining an integrated ranking for the answer based on the weighted ranking of each of the one or more human evaluators (mental process - obtaining an integrated ranking based on the weighted rankings may be performed mentally by a user analyzing the weighted rankings and accordingly using judgment/evaluation to obtain an integrated ranking based on said analysis)
determining the cumulative ranking of the answer based on the existing ranking and the integrated ranking for the answer (mental process - determining the cumulative ranking based on the existing ranking and the integrated ranking may be performed mentally by a user analyzing the existing and integrated rankings and accordingly using judgment/evaluation to determine the cumulative ranking based on said analysis)
Step 2A Prong 2 & Step 2B:
retrieving an existing ranking for the answer for the question (adding insignificant extra-solution activity to the judicial exception, mere data retrieval/gathering of the existing ranking that is then analyzed – see MPEP 2106.05(g); and, retrieving data is a well-understood, routine, and conventional function when claimed in a merely generic manner, as it is here, supporting a conclusion under Berkheimer – see MPEP 2106.05(d))
accessing the updated fidelity metric for each of the one or more human evaluators (adding insignificant extra-solution activity to the judicial exception, mere data retrieval/gathering of the fidelity metric that is then used in the analysis – see MPEP 2106.05(g); and, accessing data is a well-understood, routine, and conventional function when claimed in a merely generic manner, as it is here, supporting a conclusion under Berkheimer – see MPEP 2106.05(d))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 8. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 12:
Step 2A Prong 1: See the rejection of Claim 11 above, which Claim 12 depends on.
modifying the existing fidelity attribute based on the cumulative ranking of the answer (mental process - modifying the existing fidelity attribute based on the cumulative ranking may be performed mentally by a user analyzing the cumulative ranking and accordingly using judgment/evaluation to adjust the fidelity attribute based on said analysis)
generating an updated fidelity attribute based on the modified fidelity attribute for the machine expert (mental process - generating an updated fidelity attribute based on the modified fidelity attribute may be performed mentally or using pen and paper by a user analyzing the modified fidelity attribute and accordingly using judgment/evaluation to generate the updated fidelity attribute based on said analysis)
Step 2A Prong 2 & Step 2B:
retrieving an existing fidelity attribute associated with the machine expert (adding insignificant extra-solution activity to the judicial exception, mere data retrieval/gathering of the existing ranking that is then analyzed – see MPEP 2106.05(g); and, retrieving data is a well-understood, routine, and conventional function when claimed in a merely generic manner, as it is here, supporting a conclusion under Berkheimer – see MPEP 2106.05(d))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 11. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 13:
Step 2A Prong 1: See the rejection of Claim 8 above, which Claim 13 depends on.
Step 2A Prong 2 & Step 2B:
wherein the feedback further includes at least one of an alternative answer in place of the answer provided by one of the one or more human evaluators, an alternative reference that supports the alternative answer, and an alternative source to access the alternative reference (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the feedback further includes at least one of an alternative answer in place of the answer provided by one of the one or more human evaluators, an alternative reference that supports the alternative answer, and an alternative source to access the alternative reference does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 8. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 14:
Step 2A Prong 1: See the rejection of Claim 13 above, which Claim 14 depends on.
extracting, from the feedback, information related to the alternative reference and the alternative source (mental process - extracting information related to the alternative reference and the alternative source from the feedback may be performed mentally by a user reading/analyzing the feedback and accordingly using judgment/evaluation to identify and extract the relevant information based on said analysis)
modifying an archive storing references from different sources based on the alternative reference and the alternative source (mental process - modifying an archive storing references based on the alternative reference and the alternative source may be performed mentally or using pen and paper by a user analyzing the alternative reference and source and accordingly using judgment/evaluation to update the stored references based on said analysis)
Step 2A Prong 2 & Step 2B:
Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 15:
Step 1: Claim 15 is a system type claim. Therefore, Claims 15 - 20 fall within one of the four statutory categories (i.e., process, machine, manufacture, or composition of matter).
2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance
of the limitation in the mind but for the recitation of generic computer components, then it falls within
the “Mental Processes” grouping of abstract ideas. If a claim limitation, under its broadest reasonable
interpretation, covers performance of the limitation by mathematical calculation but for the recitation
of generic computer components, then it falls within the “Mathematical Concepts” grouping of abstract
ideas.
updating a fidelity metric associated with each of the one or more human evaluators based on the evaluation (mental process – updating a fidelity metric based on an evaluation may be performed mentally by a user observing/analyzing the received evaluation and accordingly using judgment/evaluation to adjust a value associated with each evaluator based on said analysis)
determining a cumulative ranking of the answer with respect to the question according to the evaluation and the updated fidelity metric of each of the one or more human evaluators (mental process - determining a cumulative ranking may be performed mentally by a user reading/analyzing the evaluations and the fidelity metrics and accordingly using judgment/evaluation to arrive at a ranking based on said analysis)
updating a fidelity attribute associated with the machine expert based on the cumulative ranking, wherein the fidelity attribute is indicative of an ability of the machine expert in answering questions in the subject matter (mental process - updating a fidelity attribute based on the cumulative ranking may be performed mentally by a user analyzing the cumulative ranking and accordingly using judgment/evaluation to adjust a value indicative of the machine expert's ability based on said analysis)
generating feedback based on the answer, the question, the cumulative ranking of the answer with respect to the question, and the updated fidelity attribute of the machine expert (mental process - generating feedback may be performed mentally or using pen and paper by a user reviewing/analyzing the answer, question, ranking, and fidelity attribute and accordingly using judgment/evaluation to formulate/write feedback based on said analysis)
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
[…] processor […] (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
receiving, from one or more human evaluators, evaluation directed to an answer […] (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g))
[…] automatically generated by a machine expert in a question & answer (Q&A) system in response to a question related to a subject matter based on a reference from a source (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine expert (generative AI) to generate answers without significantly more)
sending the feedback to the Q&A system for adapting the Q&A system (Adding insignificant extra-solution activity to the judicial exception - see MPEP 2106.05(g))
Step 2B: The claim does not include additional elements considered individually and in combination that
are sufficient to amount to significantly more than the judicial exception.
[…] processor […] (recited at a high-level of generality (i.e., a generic processor, computer-readable storage medium, a communication interface, a user interface and memory) such that it amounts to no more than mere instructions to apply the exception using generic computer components)
receiving, from one or more human evaluators, evaluation directed to an answer […] (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer)
[…] automatically generated by a machine expert in a question & answer (Q&A) system in response to a question related to a subject matter based on a reference from a source (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine expert (generative AI) to generate answers without significantly more)
sending the feedback to the Q&A system for adapting the Q&A system (MPEP 2106.05(d)(II) indicates that merely “Receiving or transmitting data over a network” is a well-understood, routine, conventional function when it is claimed in a merely generic manner (as it is in the present claim). Thereby, a conclusion that the claimed limitation is well-understood, routine, conventional activity is supported under Berkheimer)
For the reasons above, Claim 15 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claims 16 - 20. The additional limitations of the dependent claims are addressed below.
Regarding Claim 16:
Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 16 depends on.
Step 2A Prong 2 & Step 2B:
wherein the Q&A system includes a plurality of machine experts for automatically generating answers to questions […] (Adding the words “apply it” (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea - see MPEP 2106.05(f) – Examiner’s note: high level recitation of using a machine expert (generative AI) to generate answers without significantly more)
[…] wherein, for each question asked, at least some of the plurality of machine experts are selected for providing an answer to the question and the selection is based, at least partially, on the fidelity attribute associated with each of the plurality of machine experts space (Field of Use – limitations that amount to merely indicating a field of use or technological environment in which to apply a judicial exception does not amount to significantly more than the exception itself, and cannot integrate a judicial exception into a practical application; in this case specifying that the plurality of machine experts are selected for providing an answer to the question and the selection is based, at least partially, on the fidelity attribute associated with each of the plurality of machine experts space does not integrate the exception into a practical application nor amount to significantly more – See MPEP 2106.05(h))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the abstract idea into practical application because it does not impose any meaningful limits on practicing the abstract idea, as discussed above in the rejection of claim 15. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 17:
Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 17 depends on.
identifying a ranking for the answer provided by the human evaluator from the evaluation (mental process - identifying a ranking for the answer from the evaluation may be performed mentally by a user reading/analyzing the evaluation and accordingly using judgment/evaluation to identify the ranking based on said analysis)
determining a number of other rankings from remainder of the one or more human evaluators that are consistent with the ranking (mental process - determining a number of other rankings that are consistent with the ranking may be performed mentally or using pen and paper by a user reviewing/analyzing the rankings from the other evaluators and accordingly using judgment/evaluation to count those consistent with the ranking based on said analysis)
determining a parameter based on the number of rankings from others (mental process - determining a parameter based on the number of rankings from others may be performed mentally by a user analyzing the number of consistent rankings and accordingly using judgment/evaluation to determine a parameter based on said analysis)
updating an existing fidelity metric associated with the human evaluator based on the parameter to generate the updated fidelity metric for the human evaluator (mental process - updating an existing fidelity metric based on the parameter may be performed mentally or using pen and paper by a user analyzing the parameter and accordingly using judgment/evaluation to adjust the fidelity metric associated with the evaluator based on said analysis)
Step 2A Prong 2 & Step 2B:
Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 18:
Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 18 depends on.
identifying a ranking from the evaluation from each of the one or more human evaluators (mental process - identifying a ranking from the evaluation from each evaluator may be performed mentally by a user reading/analyzing each evaluation and accordingly using judgment/evaluation to identify the ranking based on said analysis)
weighing the ranking of each of the one or more human evaluators based on the updated fidelity metric thereof to generate a weighted ranking for the human evaluator (mental process - weighing the ranking of each evaluator based on the updated fidelity metric may be performed mentally by a user analyzing the ranking and the fidelity metric and accordingly using judgment/evaluation to generate a weighted ranking based on said analysis)
obtaining an integrated ranking for the answer based on the weighted ranking of each of the one or more human evaluators (mental process - obtaining an integrated ranking based on the weighted rankings may be performed mentally by a user analyzing the weighted rankings and accordingly using judgment/evaluation to obtain an integrated ranking based on said analysis)
determining the cumulative ranking of the answer based on the existing ranking and the integrated ranking for the answer (mental process - determining the cumulative ranking based on the existing ranking and the integrated ranking may be performed mentally by a user analyzing the existing and integrated rankings and accordingly using judgment/evaluation to determine the cumulative ranking based on said analysis)
Step 2A Prong 2 & Step 2B:
retrieving an existing ranking for the answer for the question (adding insignificant extra-solution activity to the judicial exception, mere data retrieval/gathering of the existing ranking that is then analyzed – see MPEP 2106.05(g); and, retrieving data is a well-understood, routine, and conventional function when claimed in a merely generic manner, as it is here, supporting a conclusion under Berkheimer – see MPEP 2106.05(d))
accessing the updated fidelity metric for each of the one or more human evaluators (adding insignificant extra-solution activity to the judicial exception, mere data retrieval/gathering of the fidelity metric that is then used in the analysis – see MPEP 2106.05(g); and, accessing data is a well-understood, routine, and conventional function when claimed in a merely generic manner, as it is here, supporting a conclusion under Berkheimer – see MPEP 2106.05(d))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 15. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 19:
Step 2A Prong 1: See the rejection of Claim 18 above, which Claim 19 depends on.
modifying the existing fidelity attribute based on the cumulative ranking of the answer (mental process - modifying the existing fidelity attribute based on the cumulative ranking may be performed mentally by a user analyzing the cumulative ranking and accordingly using judgment/evaluation to adjust the fidelity attribute based on said analysis)
generating an updated fidelity attribute based on the modified fidelity attribute for the machine expert (mental process - generating an updated fidelity attribute based on the modified fidelity attribute may be performed mentally or using pen and paper by a user analyzing the modified fidelity attribute and accordingly using judgment/evaluation to generate the updated fidelity attribute based on said analysis)
Step 2A Prong 2 & Step 2B:
retrieving an existing fidelity attribute associated with the machine expert (adding insignificant extra-solution activity to the judicial exception, mere data retrieval/gathering of the existing ranking that is then analyzed – see MPEP 2106.05(g); and, retrieving data is a well-understood, routine, and conventional function when claimed in a merely generic manner, as it is here, supporting a conclusion under Berkheimer – see MPEP 2106.05(d))
Accordingly, under Step 2A Prong 2 and Step 2B, this additional element does not integrate the
abstract idea into practical application because it does not impose any meaningful limits on practicing
the abstract idea, as discussed above in the rejection of claim 18. The claim does not include additional
elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Regarding Claim 20:
Step 2A Prong 1: See the rejection of Claim 15 above, which Claim 20 depends on.
extracting, from the feedback, information related to at least one of an alternative answer in place of the answer provided by one of the one or more human evaluators, an alternative reference relied on to derive the alterative answer, and an alternative source to access the alternative reference (mental process - extracting information related to at least one of an alternative answer in place of the answer provided by one of the one or more human evaluators may be performed mentally by a user reading/analyzing the feedback and accordingly using judgment/evaluation to identify and extract the relevant information based on said analysis)
modifying an archive storing references from different sources based on the alternative reference and the alternative source (mental process - modifying an archive storing references based on the alternative reference and the alternative source may be performed mentally or using pen and paper by a user analyzing the alternative reference and source and accordingly using judgment/evaluation to update the stored references based on said analysis)
Step 2A Prong 2 & Step 2B:
Accordingly, under Step 2A Prong 2 and Step 2B, there are no additional elements that integrate the abstract idea into practical application. The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 - 20 are rejected under 35 U.S.C. 103 as being unpatentable over Johnson et al. (hereinafter Johnson) (US 20140280087) in view of Jamaludeen et al. (hereinafter Jamaludeen, a non-patent literature reference titled “Assessing the reliability of crowdsourced labels via Twitter”), and in further view of Wang et al. (hereinafter Wang) (US 20180096283).
Regarding Claim 1, Johnson teaches:
receiving, from one or more human evaluators, evaluation directed to an answer automatically generated by a machine expert in a question & answer (Q&A) system in response to a question related to a subject matter based on a reference from a source (Johnson, Par. [0093], “The user feedback rating engine 720 provides logic for presenting a listing of candidate answers generated by the QA system 710 and provided in the results 715, as well as their confidence measures, answer source information, e.g., corpus, document, etc., and the like, via a user interface 725. The user feedback rating engine 720 further provides the user interface 725 with elements and/or fields through which user input is received for rating the quality of the candidate answers and/or answer source information presented in the results 715”, & Par. [0071], “The QA system may comprise multiple engines or modules comprising logic for performing various operations for processing an input question in a natural language, searching a corpus of information for generating candidate answers to the input question, ranking or scoring the candidate answers, and performing a final merging of the scored or ranked candidate answers to generate a single ultimate answer to the input question. Thus, the QA system may comprise engines/modules for performing question analysis, content analysis of documents in a corpus of information, primary search, candidate answer generation, candidate answer scoring/ranking, and final merging of candidate answers”, & Par. [0074], “For a natural language question raised by a user, question parsing and focus detecting are performed in the question processing module 501, which generates queries for the question”, & Par. [0074], “Afterwards, the answer processing module 505 performs candidate identification and answer ranking on the candidate answers generated by the document/passage retrieval module 503, and finally formulates an answer to the raised natural language question, so as to output a brief answer to the user in natural language”, thus Johnson teaches a QA system having multiple engines/modules configured to process an input question, search a corpus of information, generate candidate answers, and formulate an answer to the question, and further teaches receiving user input that rates the quality of the generated candidate answers and associated answer-source information. Johnson’s candidate-answer-generation engines/modules correspond to a machine expert because the engines/modules automatically generate candidate answers in response to an input question. Johnson’s candidate answers correspond to an answer, the user providing a rating corresponds to a human evaluator, the user rating corresponds to evaluation, and Johnson’s corpus/document and answer source information correspond to a reference from a source)
[…] the one or more human evaluators based on the evaluation (Johnson, Par. [0031], “human operators, analysts, or the like, are able to provide feedback ratings of candidate answers to posed questions during the training process and these feedback ratings are automatically processed in addition to the processing of the generated candidate answers to determine confidence measures for the generated candidate answers and refine a training model for performing machine learning during the training of the QA systems”, thus […] the one or more human evaluators based on the evaluation is disclosed, because Johnson teaches that human operators or analysts provide feedback ratings of candidate answers and that those feedback ratings are automatically processed to determine confidence measures and refine the training model. Johnson’s human operators or analysts correspond to the one or more human evaluators, and the feedback ratings correspond to the evaluation)
determining a cumulative ranking of the answer with respect to the question according to the evaluation […] of each of the one or more human evaluators (Johnson, Par. [0098], “the user feedback ratings of the candidate answers may further be used to modify the logic or configuration parameters so as to increase the confidence measures of the answers the users found to be most useful or having the highest quality while reducing those that have a relatively lower user feedback rating”, thus determining a cumulative ranking of the answer with respect to the question according to the evaluation […] of each of the one or more human evaluators is disclosed, because Johnson teaches using user feedback ratings of candidate answers to increase or decrease the confidence measures associated with the answers according to the users’ assessment of their usefulness or quality. Johnson’s user feedback ratings correspond to the evaluation, and the confidence measures associated with the candidate answers correspond to the ranking of the answer)
generating feedback based on the answer, the question, the cumulative ranking of the answer with respect to the question, […] the machine expert (Johnson, Par. [0073], “The illustrative embodiments provide mechanisms for automatically identifying areas of improvement, provide mechanisms for generating recommendations for such improvement, and provide mechanisms for actually improving the operation of a QA system with regard to the accuracy or confidence measures of candidate answers generated by the QA system. As mentioned previously, the illustrative embodiments provide mechanisms for answering questions posed regarding the results generated by the training of the QA system, mechanisms for allowing users to input ratings of the candidate answers which may then be processed by the QA system to determine improvements to the model or logic used by the QA system”, & Par. [0098], “the user feedback ratings of the candidate answers may further be used to modify the logic or configuration parameters so as to increase the confidence measures of the answers the users found to be most useful or having the highest quality while reducing those that have a relatively lower user feedback rating”, & Par. [0099], “the user poses the training question "Who was the man that crossed the Potomac?" There may be multiple candidate answers that may be returned by the QA system 710 in the results 715 with separate confidence measures for each of these candidate answers”, thus generating feedback based on the answer, the question, the cumulative ranking of the answer with respect to the question, […] the machine expert is disclosed, because Johnson teaches generating recommendations for improving the QA system based on candidate answers, the corresponding questions, user ratings of the candidate answers, and confidence measures associated with those answers. Johnson’s generated recommendations correspond to feedback, the candidate answers correspond to the answer, the posed training question corresponds to the question, and the confidence measures adjusted according to the user feedback ratings correspond to the cumulative ranking of the answer with respect to the question)
sending the feedback to the Q&A system for adapting the Q&A system (Johnson, Par. [0073], “mechanisms for allowing users to input ratings of the candidate answers which may then be processed by the QA system to determine improvements to the model or logic used by the QA system”, & Par. [0096], “the user rating feedback information may be used to refine the training performed by the training engine 740 based on the identified differences between the results 715 and the correct answers specified in the ground truth data storage 730. The modifications to the QA system 710 based on user rating feedback may be implemented, for example, in various ones of the modules/engines of the QA system 710”, & Par. [0101], “Thus, the user feedback may be used to adjust the training of the QA system 710”, thus sending the feedback to the Q&A system for adapting the Q&A system is disclosed, because Johnson teaches providing user ratings of candidate answers to the QA system, processing those ratings to determine improvements to the QA system’s model or logic, and using the feedback to refine training and modify modules or engines of the QA system. Johnson’s user rating feedback corresponds to the feedback, and refining the training and modifying the model, logic, or modules of the QA system corresponds to adapting the Q&A system)
Johnson does not explicitly teach updating a fidelity metric associated with each of […human evaluators…], […] and the updated fidelity metric […], updating a fidelity metric associated with each of […human evaluators…], updating a fidelity attribute associated with […the machine expert based on the cumulative ranking…], wherein the fidelity attribute is indicative of an ability of […the machine expert…] in answering questions in […the subject matter…], and […] and the updated fidelity attribute of […].
However, Jamaludeen teaches:
updating a fidelity metric associated with each of […human evaluators…] (Jamaludeen, Page 4 – Section 3.3, “To distinguish reliable annotators from unreliable ones, we introduce the concept of reliability score: for each tweet y ∈W annotated by x, we set agreement(x,y)= 0, if vote(x,y) = inferred label(y) 1, if vote(x,y) = inferred label(y) (1) Then,we define the reliability score of annotator x over topic j asRSx,j,whereRSx,j = y∈W ∧TPy,j=0 agreement(x,y)”, & Page 5 – Section 3.4, “We start processing with top-1 instance in listW,weinferthelabelofthisinstanceusingtheinitialreliabilityscores, update the reliability scores for topics comprised in the top-1 instance according to its inferred label, then move to infer the label of top-2 instance employing the up dated reliability scores, reupdate again the reliability socres accordingly and so on”, thus updating a fidelity metric associated with each of […human evaluators…] is disclosed, because Jamaludeen teaches assigning a reliability score to each annotator and incrementally updating that reliability score according to the annotator’s vote relative to the inferred label. Jamaludeen’s annotator corresponds to a human evaluator, and the annotator’s reliability score corresponds to a fidelity metric associated with the human evaluator)
[…] and the updated fidelity metric […] (Jamaludeen, Page 4 – Section 3.4, “The votes are weighted with the annotators’ topic-based reliability scores”, & Page 5 – Section 3.4, “Each vote is weighted with the sum of the annotator tweet-related topic reliability scores. The weights are aggregated for annotators who provided identical votes by summing them up”, thus […] and the updated fidelity metric […] is disclosed, because Jamaludeen teaches weighting each annotator’s vote using the annotator’s topic-based reliability score and then aggregating the weighted votes. Jamaludeen’s topic-based reliability score corresponds to the updated fidelity metric, and using that score to weight each annotator’s vote corresponds to determining the result according to the updated fidelity metric)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Johnson with Jamaludeen by incorporating Jamaludeen’s annotator reliability scoring and weighted voting technique into Johnson’s Q&A user feedback system. Johnson teaches receiving feedback ratings from human operators or analysts and processing those ratings to determine confidence measures for generated candidate answers and refine the training of the Q&A system. Jamaludeen teaches assigning and updating reliability scores for annotators and weighting their evaluations according to those reliability scores. Therefore, a POSITA would have been motivated to use Jamaludeen’s reliability scores to weight the feedback ratings received from Johnson’s human operators so that feedback from more reliable evaluators would have a greater effect on the resulting answer ranking, while feedback from less reliable evaluators would have a reduced effect, thereby improving the accuracy and robustness of Johnson’s feedback processing, reducing the influence of unreliable human evaluations, and improving the reliability of the resulting answer evaluation and Q&A system performance (Jamaludeen, Page 11 – Section 5.6, “As a result, the more anno tations the annotator delivers, the model’s capacity of estimating the annotator’s re liability scores improves, thus the labels inference enhances. The model was also robust across different percentages of unreliable annotators and performed better than the Kappa Weighted Voting approach”)
Johnson combined with Jamaludeen does not explicitly teach updating a fidelity attribute associated with […the machine expert based on the cumulative ranking…], wherein the fidelity attribute is indicative of an ability of […the machine expert…] in answering questions in […the subject matter…], and […] and the updated fidelity attribute of […].
However, Wang teaches:
updating a fidelity attribute associated with […the machine expert based on the cumulative ranking…], wherein the fidelity attribute is indicative of an ability of […the machine expert…] in answering questions in […the subject matter…] (Wang, Par. [0101], “Assistant module 222 may determine a score of the experience, and feed the determined score back into ranking. For instance, assistant module 222 may modify the agent-quality score of the agent that fulfilled the request based on the user’s feedback about the fulfillment”, & Par. [0114], “the agent directory may collect user reviews and ratings. The collected user reviews and ratings may be used to modify the agent quality scores. As one example, when an agent receives positive reviews and/or ratings, agent accuracy module 331 may increase the agent’s agent quality score”, & Par. [0057], “local assistant module 122A may select a 3P agent based on rankings. For instance, local assistant module 122A may select a 3P agent with the highest ranking to satisfy the utterance”, thus updating a fidelity attribute associated with […the machine expert based on the cumulative ranking…], wherein the fidelity attribute is indicative of an ability of […the machine expert…] in answering questions in […the subject matter…] is disclosed, because Wang teaches determining an experience score from user feedback, feeding that score back into the ranking, and modifying the agent quality score of the particular agent that fulfilled the request. Wang further teaches increasing or otherwise modifying an agent’s quality score based on user reviews and ratings and selecting an agent based on its ranking for satisfying an utterance. Wang’s agent corresponds to a machine expert, the agent-quality score corresponds to a fidelity attribute, and the ranking used to select an agent for satisfying an utterance corresponds to an indication of the agent’s ability with respect to the subject matter)
[…] and the updated fidelity attribute of […] (Wang, Par. [0101], “assistant module 222 may modify the agent-quality score of the agent that fulfilled the request based on the user’s feedback about the fulfillment”, & Par. [0114], “The collected user reviews and ratings may be used to modify the agent quality scores”, thus […] and the updated fidelity attribute of […] is disclosed, because Wang teaches modifying an existing agent quality score based on user feedback and further modifying agent quality scores based on collected user reviews and ratings. Wang’s modified agent quality score corresponds to the updated fidelity attribute)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Johnson and Jamaludeen with Wang by incorporating Wang’s agent quality scoring and agent selection technique into the combined Q&A user feedback system. Johnson and Jamaludeen teach processing human evaluations while accounting for evaluator reliability to provide a more reliable evaluation of generated answers. Wang teaches evaluating and ranking individual agents so that an appropriate agent may be selected for a particular task. Therefore, a POSITA would have been motivated to use the reliability weighted evaluation provided by Johnson and Jamaludeen to update the quality score of the agent that generated the evaluated answer and use that score in selecting an appropriate agent for subsequent questions, thereby improving user machine interaction, increasing the likelihood that an appropriate agent is selected for a particular task, and reducing unnecessary consumption of computing resources and power by less appropriate agents (Wang, Par. [0004], “a computational assistant may initially process the utterance and select an appropriate agent to respond to the utterance. That is, the most appropriate agent may be selected for any given task received in an utterance. This provides an adaptive interface which improves user-machine interaction. Having a single assistant initially process the utterance before an appropriate agent responds to the utterance prevents other (less appropriate) agents from wasting computing resources or consuming power, processing the utterance. Accordingly, the techniques of this disclosure may enable a system to use less power and system resources than other systems that enable multiple agents to simultaneously process received utterances”)
Regarding Claim 2, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 1 as cited above and Johnson further teaches:
wherein the Q&A system includes a plurality of machine experts for automatically generating answers to questions […] at least some of the plurality of machine experts […] associated with each of the plurality of machine experts (Johnson, Par. [0071], “The QA system may comprise multiple engines or modules comprising logic for performing various operations for processing an input question in a natural language, searching a corpus of information for generating candidate answers to the input question, ranking or scoring the candidate answers, and performing a final merging of the scored or ranked candidate answers to generate a single ultimate answer to the input question,” thus wherein the Q&A system includes a plurality of machine experts for automatically generating answers to questions […] at least some of the plurality of machine experts […] associated with each of the plurality of machine experts is disclosed, because Johnson teaches a Q&A system having multiple engines or modules that process input questions, generate candidate answers, rank or score those candidate answers, and merge the ranked or scored answers to generate an ultimate answer. Johnson’s multiple engines or modules correspond to the plurality of machine experts, and the candidate answer generation performed by the engines or modules corresponds to automatically generating answers to questions)
Wang further teaches:
[…], wherein, for each question asked, […] are selected for providing an answer to the question and the selection is based, at least partially, on the fidelity attribute […] (Wang, Par. [0053], “local assistant module 122 may select an agent from a plurality of agents to satisfy the utterance”, & Par. [0057], “local assistant module 122A may select a 3P agent with the highest ranking to satisfy the utterance. In some examples, such as where there is a tie in the rankings and/or if the ranking of the 3P agent with the highest ranking is less than a ranking threshold, local assistant module 122A may solicit user input to select a 3P language agent to satisfy the utterance”, & Par. [0094], “Agent selection module 227 may rank the identified agent documents (e.g., based on a capability level to satisfy the utterance). For instance, agent selection module 227 may multiply a text-match score with an agent-quality score,” thus […], wherein, for each question asked, […] are selected for providing an answer to the question and the selection is based, at least partially, on the fidelity attribute […] is disclosed, because Wang teaches selecting an agent from a plurality of agents to satisfy an utterance, selecting a highest ranked agent to satisfy the utterance, and ranking candidate agents based on a capability level to satisfy the utterance using an agent-quality score. Wang’s selection of an agent to satisfy the utterance corresponds to selecting for providing an answer to the question, and Wang’s use of the agent quality score in ranking the agents corresponds to basing the selection, at least partially, on the fidelity attribute)
Regarding Claim 3, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 1 as cited above and Johnson further teaches:
identifying a ranking for the answer provided by the human evaluator from the evaluation (Johnson, Par. [0100], “The user may specify, for example, that the candidate answer "George Washington" is the most useful answer, or the answer having the highest quality, relative to the other candidate answers, and that the answers "General Washington" and "Washington" have relatively lower and lowest user ratings”, thus identifying a ranking for the answer provided by the human evaluator from the evaluation is disclosed, because Johnson teaches receiving a user evaluation that identifies one candidate answer as having the highest quality relative to other candidate answers and assigns relatively lower ratings to the remaining answers. Johnson’s relative user ratings correspond to the ranking for the answer identified from the human evaluator’s evaluation)
Jamaludeen further teaches:
determining a number of other rankings from remainder of […the one or more human evaluators…] that are consistent with […the ranking…] (Jamaludeen, Page 4 – Section 3.2, “For tweet y and class label c, let votes(y,c) be the number of annotators who as signed c to y. We assign each tweet to the class according to the majority voting, i.e. mvlabel(y) = argmaxc∈Cvotes(y,c).We use this number also to assign a rank to y” & “The rank reflects the agreement of annotators concerning the selected class label according to the majority voting labelling. We consider consensus as indicator of how much the class label of the instance can be trusted, and process high-ranked instances before low-ranked instances when computing annotator reliability”, thus determining a number of other rankings from remainder of […the one or more human evaluators…] that are consistent with […the ranking…] is disclosed, because Jamaludeen teaches counting the number of annotators that assign a particular class label and using that number to determine a majority label and corresponding rank reflecting agreement among the annotators. Jamaludeen’s annotators correspond to […the one or more human evaluators…], the annotator votes correspond to the rankings, and the number of annotators assigning the same class label corresponds to the number of other rankings that are consistent with […the ranking…])
determining a parameter based on the number of rankings from others (Jamaludeen, Page 4 – Section 3.2, “We assign each tweet to the class according to the majority voting, i.e. mvlabel(y) = argmaxc∈Cvotes(y,c).We use this number also to assign a rank to y”, & Page 4 – Section 3.3, “To distinguish reliable annotators from unreliable ones, we introduce the concept of reliability score: for each tweet y ∈W annotated by x, we set agreement(x,y)= 0, if vote(x,y)=inferredlabel(y) 1, if vote(x,y)=inferredlabel(y)”, thus determining a parameter based on the number of rankings from others is disclosed, because Jamaludeen teaches determining an inferred label from the number of annotator votes and then determining an agreement value according to whether an annotator’s vote is consistent with the inferred label. Jamaludeen’s agreement value corresponds to the parameter, and because the inferred label is determined from the number of votes from the other annotators, the agreement value is based on the number of rankings from others)
updating an existing fidelity metric associated with […the human evaluator…] based on the parameter to generate the updated fidelity metric for […the human evaluator…] (Jamaludeen, Page 4 – Section 3.4, “The tweets y ∈ W are processed incrementally and the reliability scores are updated simultaneously”, & Pge 5 – Section 3.4, “For each annotator who gave a vote identical to the inferred label, increment the tweet-related topic reliability scores by 1 as follows: RSx,j,t ←RSx,j,t−1+1”, thus updating an existing fidelity metric associated with […the human evaluator…] based on the parameter to generate the updated fidelity metric for […the human evaluator…] is disclosed, because Jamaludeen teaches incrementally updating annotator reliability scores and increasing the reliability score of an annotator when the annotator’s vote is identical to the inferred label. Jamaludeen’s annotator corresponds to […the human evaluator…], the existing reliability score corresponds to the existing fidelity metric, the determination that the annotator’s vote matches the inferred label corresponds to the parameter, and the incremented reliability score corresponds to the updated fidelity metric for […the human evaluator…])
Regarding Claim 4, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 1 as cited above and Johnson further teaches:
retrieving an existing ranking for the answer for the question (Johnson, Par. [0076], “The training run results query engine 510 provides logic for allowing users to use the QA system 500 to answer questions about the results generated by the QA system 500 of a previous run, or execution, of the QA system on a previously submitted question”, & Par. [0077], “results from inputting a training question to the QA system 500 are generated by the QA system 500 processing a training set of data to generate a set of candidate answers with corresponding confidence measures”, thus retrieving an existing ranking for the answer for the question is disclosed, because Johnson teaches accessing results from a previous execution of the QA system for a previously submitted question, wherein the results include candidate answers having corresponding confidence measures. Johnson’s previously generated confidence measure associated with a candidate answer corresponds to the existing ranking for the answer for the question)
determining the cumulative ranking of the answer based on the existing ranking […] (Johnson, Par. [0098], “the user feedback ratings of the candidate answers may further be used to modify the logic or configuration parameters so as to increase the confidence measures of the answers the users found to be most useful or having the highest quality while reducing those that have a relatively lower user feedback rating”), thus determining the cumulative ranking of the answer based on the existing ranking […] is disclosed, because Johnson teaches modifying an existing confidence measure associated with an answer based on user feedback ratings. Johnson’s existing confidence measure corresponds to the existing ranking, and the modified confidence measure corresponds to the cumulative ranking)
Jamaludeen further teaches:
accessing the updated fidelity metric for each of […the one or more human evaluators…] (Jamaludeen, Page 5 – Section 3.4, “we infer the label of this instance using the initial reliability scores, update the reliability scores for topics comprised in the top-1 instance according to its inferred label, then move to infer the label of top-2 instance employing the up dated reliability scores”, thus accessing the updated fidelity metric for each of […the one or more human evaluators…] is disclosed, because Jamaludeen teaches using previously updated annotator reliability scores when processing a subsequent instance. Jamaludeen’s annotators correspond to […the one or more human evaluators…], and the updated annotator reliability scores correspond to the updated fidelity metrics associated with those evaluators)
identifying a ranking from the evaluation from each of [… the one or more human evaluators...] (Jamaludeen, Page 5 – Section 3.4, “Each vote is weighted with the sum of the annotator tweet related topic reliability scores”, & “The weights are aggregated for annotators who provided identical votes by summing them up”, thus identifying a ranking from the evaluation from each of […the one or more human evaluators…] is disclosed, because Jamaludeen teaches identifying the individual vote provided by each annotator before weighting and aggregating the votes. Jamaludeen’s annotators correspond to […the one or more human evaluators…], and each annotator’s vote corresponds to a ranking identified from that evaluator’s evaluation)
weighing the ranking […of each of the one or more human evaluators…] based on the updated fidelity metric thereof to generate a weighted ranking for […the human evaluator…] (Jamaludeen, Page 5 – Section 3.4, “Each vote is weighted with the sum of the annotator tweet related topic reliability scores”, thus weighing the ranking […of each of the one or more human evaluators…] based on the updated fidelity metric thereof to generate a weighted ranking for […the human evaluator…] is disclosed, because Jamaludeen teaches weighting each annotator’s vote using the annotator’s reliability scores. Jamaludeen’s annotator corresponds to […the one or more human evaluators…], the annotator’s vote corresponds to the ranking, the annotator’s reliability score corresponds to the updated fidelity metric, and the resulting weighted vote corresponds to the weighted ranking for […the human evaluator…])
obtaining an integrated ranking for the answer based on the weighted ranking of […each of the one or more human evaluators…] (Jamaludeen, Page 5 – Section 3.4, “The weights are aggregated for annotators who provided identical votes by summing them up”, & “We select the class label that collected the highest weight as the label for the tweet”, thus obtaining an integrated ranking for the answer based on the weighted ranking of […each of the one or more human evaluators…] is disclosed, because Jamaludeen teaches aggregating the weighted votes of multiple annotators and determining the resulting outcome based on the accumulated weights. Jamaludeen’s weighted annotator votes correspond to the weighted rankings of […each of the one or more human evaluators…], and the aggregated weight used to determine the resulting label corresponds to the integrated ranking for the answer)
[…] and the integrated ranking for […the answer…] (Jamaludeen, Page 5 - Section 3.4, “The weights are aggregated for annotators who provided identical votes by summing them up”, & “We select the class label that collected the highest weight as the label for the tweet”), thus […] and the integrated ranking for […the answer…] is disclosed, because Jamaludeen teaches aggregating the weighted votes of multiple annotators and determining the resulting outcome based on the accumulated weights. Jamaludeen’s aggregated weighted votes correspond to the integrated ranking for […the answer…])
Regarding Claim 5, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 4 as cited above and Wang further teaches:
retrieving an existing fidelity attribute associated with […the machine expert…] (Wang, Par. [0047], “agent indices 124 may store information related to the use and/or the performance of the available agents. For instance, agent indices 124 may include an agent-quality score for each available agent”, thus retrieving an existing fidelity attribute associated with […the machine expert…] is disclosed, because Wang teaches storing an agent-quality score associated with each available agent. Wang’s agent corresponds to […the machine expert…], and the stored agent-quality score corresponds to the existing fidelity attribute associated with […the machine expert…])
modifying the existing fidelity attribute based on […the cumulative ranking of the answer…] (Wang, Par. [0101], “assistant module 222 may modify the agent-quality score of the agent that fulfilled the request based on the user's feedback about the fulfillment”, & Par. [0114], “the agent directory may collect user reviews and ratings. The collected user reviews and ratings may be used to modify the agent quality scores”, thus modifying the existing fidelity attribute based on […the cumulative ranking of the answer…] is disclosed, because Wang teaches modifying an existing agent-quality score based on feedback and ratings concerning the performance of the agent. Wang’s existing agent-quality score corresponds to the existing fidelity attribute)
generating an updated fidelity attribute based on the modified fidelity attribute for […the machine expert…] (Wang, Par. [0114], “when an agent receives positive reviews and/or ratings, agent accuracy module 331 may increase the agent's agent quality score in agent index 224 or agent index 324. As another example, when an agent receives negative reviews and/or ratings, agent accuracy module 331 may decrease the agent's agent quality score in agent index 224 or agent index 324”, thus generating an updated fidelity attribute based on the modified fidelity attribute for […the machine expert…] is disclosed, because Wang teaches modifying an existing agent quality score by increasing or decreasing the score based on received reviews and ratings, thereby producing an updated agent quality score. Wang’s updated agent quality score corresponds to the updated fidelity attribute for […the machine expert…])
Regarding Claim 6, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 1 as cited above and Johnson further teaches:
wherein the feedback further includes at least one of an alternative answer in place of the answer provided by one of the one or more human evaluators, an alternative reference that supports the alternative answer, and an alternative source to access the alternative reference (Johnson, Par. [0093], “The user feedback rating engine 720 provides logic for presenting a listing of candidate answers generated by the QA system 710 and provided in the results 715, as well as their confidence measures, answer source information, e.g., corpus, document, etc., and the like, via a user interface 725”, & Par. [0100], “the user is presented with a listing of the candidate answers and options or fields through which a user may specify a rating of each of the candidate answers, via the user feedback rating engine 720”, thus wherein the feedback further includes at least one of an alternative answer in place of the answer provided by one of the one or more human evaluators, an alternative reference that supports the alternative answer, and an alternative source to access the alternative reference is disclosed, because Johnson teaches presenting multiple candidate answers and corresponding answer source information to a user and receiving user feedback ratings concerning those candidate answers. Johnson’s candidate answers other than a particular answer correspond to an alternative answer in place of the answer, and Johnson’s corpus or document information associated with the candidate answers corresponds to an alternative reference or alternative source supporting the alternative answer)
Regarding Claim 7, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 6 as cited above and Johnson further teaches:
extracting, from the feedback, information related to the alternative reference and the alternative source (Johnson, Par. [0093], “The user feedback rating engine 720 further provides the user interface 725 with elements and/or fields through which user input is received for rating the quality of the candidate answers and/or answer source information presented in the results 715”, & Par. [0102], “the QA system 710 may identify the sources upon which the QA system 710 relies for each of the candidate answers in the results 715”, thus extracting, from the feedback, information related to the alternative reference and the alternative source is disclosed, because Johnson teaches receiving user feedback directed to candidate answers and associated answer source information and identifying the sources relied upon for the candidate answers. Johnson’s answer source information corresponds to information related to the alternative reference and the alternative source, and identifying that source information from the feedback corresponds to extracting the information related to the alternative reference and alternative source)
modifying an archive storing references from different sources based on the alternative reference and the alternative source (Johnson, Par. [0074], “the document/passage retrieval module 503 applies the queries to a corpus of information 502, such as a database of structure and/or unstructured document data, and performs document filtering and passage post-filtering in a document containing the content, e.g., keywords matching criteria of one or more of the queries, so as to generate candidate answers”, & Par. [0102], “the user feedback, obtained via the user feedback rating engine 720, may be accumulated for an answer source and may be used to modify the confidence measures associated with the answer source”, thus modifying an archive storing references from different sources based on the alternative reference and the alternative source is disclosed, because Johnson teaches using a corpus of information comprising structured or unstructured document data as the source from which candidate answers are generated, and further teaches modifying information associated with an answer source based on accumulated user feedback. Johnson’s corpus of information corresponds to the archive storing references from different sources, the documents in the corpus correspond to the references, and modifying the confidence measures associated with an answer source based on feedback corresponds to modifying information in the archive based on the alternative reference and the alternative source)
Regarding Claim 8, Johnson teaches:
receiving, from one or more human evaluators, evaluation directed to an answer automatically generated by a machine expert in a question & answer (Q&A) system in response to a question related to a subject matter based on a reference from a source (Johnson, Par. [0093], “The user feedback rating engine 720 provides logic for presenting a listing of candidate answers generated by the QA system 710 and provided in the results 715, as well as their confidence measures, answer source information, e.g., corpus, document, etc., and the like, via a user interface 725. The user feedback rating engine 720 further provides the user interface 725 with elements and/or fields through which user input is received for rating the quality of the candidate answers and/or answer source information presented in the results 715”, & Par. [0071], “The QA system may comprise multiple engines or modules comprising logic for performing various operations for processing an input question in a natural language, searching a corpus of information for generating candidate answers to the input question, ranking or scoring the candidate answers, and performing a final merging of the scored or ranked candidate answers to generate a single ultimate answer to the input question. Thus, the QA system may comprise engines/modules for performing question analysis, content analysis of documents in a corpus of information, primary search, candidate answer generation, candidate answer scoring/ranking, and final merging of candidate answers”, & Par. [0074], “For a natural language question raised by a user, question parsing and focus detecting are performed in the question processing module 501, which generates queries for the question”, & Par. [0074], “Afterwards, the answer processing module 505 performs candidate identification and answer ranking on the candidate answers generated by the document/passage retrieval module 503, and finally formulates an answer to the raised natural language question, so as to output a brief answer to the user in natural language”, thus Johnson teaches a QA system having multiple engines/modules configured to process an input question, search a corpus of information, generate candidate answers, and formulate an answer to the question, and further teaches receiving user input that rates the quality of the generated candidate answers and associated answer-source information. Johnson’s candidate-answer-generation engines/modules correspond to a machine expert because the engines/modules automatically generate candidate answers in response to an input question. Johnson’s candidate answers correspond to an answer, the user providing a rating corresponds to a human evaluator, the user rating corresponds to evaluation, and Johnson’s corpus/document and answer source information correspond to a reference from a source)
[…] the one or more human evaluators based on the evaluation (Johnson, Par. [0031], “human operators, analysts, or the like, are able to provide feedback ratings of candidate answers to posed questions during the training process and these feedback ratings are automatically processed in addition to the processing of the generated candidate answers to determine confidence measures for the generated candidate answers and refine a training model for performing machine learning during the training of the QA systems”, thus […] the one or more human evaluators based on the evaluation is disclosed, because Johnson teaches that human operators or analysts provide feedback ratings of candidate answers and that those feedback ratings are automatically processed to determine confidence measures and refine the training model. Johnson’s human operators or analysts correspond to the one or more human evaluators, and the feedback ratings correspond to the evaluation)
determining a cumulative ranking of the answer with respect to the question according to the evaluation […] of each of the one or more human evaluators (Johnson, Par. [0098], “the user feedback ratings of the candidate answers may further be used to modify the logic or configuration parameters so as to increase the confidence measures of the answers the users found to be most useful or having the highest quality while reducing those that have a relatively lower user feedback rating”, thus determining a cumulative ranking of the answer with respect to the question according to the evaluation […] of each of the one or more human evaluators is disclosed, because Johnson teaches using user feedback ratings of candidate answers to increase or decrease the confidence measures associated with the answers according to the users’ assessment of their usefulness or quality. Johnson’s user feedback ratings correspond to the evaluation, and the confidence measures associated with the candidate answers correspond to the ranking of the answer)
generating feedback based on the answer, the question, the cumulative ranking of the answer with respect to the question, […] the machine expert (Johnson, Par. [0073], “The illustrative embodiments provide mechanisms for automatically identifying areas of improvement, provide mechanisms for generating recommendations for such improvement, and provide mechanisms for actually improving the operation of a QA system with regard to the accuracy or confidence measures of candidate answers generated by the QA system. As mentioned previously, the illustrative embodiments provide mechanisms for answering questions posed regarding the results generated by the training of the QA system, mechanisms for allowing users to input ratings of the candidate answers which may then be processed by the QA system to determine improvements to the model or logic used by the QA system”, & Par. [0098], “the user feedback ratings of the candidate answers may further be used to modify the logic or configuration parameters so as to increase the confidence measures of the answers the users found to be most useful or having the highest quality while reducing those that have a relatively lower user feedback rating”, & Par. [0099], “the user poses the training question "Who was the man that crossed the Potomac?" There may be multiple candidate answers that may be returned by the QA system 710 in the results 715 with separate confidence measures for each of these candidate answers”, thus generating feedback based on the answer, the question, the cumulative ranking of the answer with respect to the question, […] the machine expert is disclosed, because Johnson teaches generating recommendations for improving the QA system based on candidate answers, the corresponding questions, user ratings of the candidate answers, and confidence measures associated with those answers. Johnson’s generated recommendations correspond to feedback, the candidate answers correspond to the answer, the posed training question corresponds to the question, and the confidence measures adjusted according to the user feedback ratings correspond to the cumulative ranking of the answer with respect to the question)
sending the feedback to the Q&A system for adapting the Q&A system (Johnson, Par. [0073], “mechanisms for allowing users to input ratings of the candidate answers which may then be processed by the QA system to determine improvements to the model or logic used by the QA system”, & Par. [0096], “the user rating feedback information may be used to refine the training performed by the training engine 740 based on the identified differences between the results 715 and the correct answers specified in the ground truth data storage 730. The modifications to the QA system 710 based on user rating feedback may be implemented, for example, in various ones of the modules/engines of the QA system 710”, & Par. [0101], “Thus, the user feedback may be used to adjust the training of the QA system 710”, thus sending the feedback to the Q&A system for adapting the Q&A system is disclosed, because Johnson teaches providing user ratings of candidate answers to the QA system, processing those ratings to determine improvements to the QA system’s model or logic, and using the feedback to refine training and modify modules or engines of the QA system. Johnson’s user rating feedback corresponds to the feedback, and refining the training and modifying the model, logic, or modules of the QA system corresponds to adapting the Q&A system)
Johnson does not explicitly teach updating a fidelity metric associated with each of […human evaluators…], […] and the updated fidelity metric […], updating a fidelity metric associated with each of […human evaluators…], updating a fidelity attribute associated with […the machine expert based on the cumulative ranking…], wherein the fidelity attribute is indicative of an ability of […the machine expert…] in answering questions in […the subject matter…], and […] and the updated fidelity attribute of […].
However, Jamaludeen teaches:
updating a fidelity metric associated with each of […human evaluators…] (Jamaludeen, Page 4 – Section 3.3, “To distinguish reliable annotators from unreliable ones, we introduce the concept of reliability score: for each tweet y ∈W annotated by x, we set agreement(x,y)= 0, if vote(x,y) = inferred label(y) 1, if vote(x,y) = inferred label(y) (1) Then,we define the reliability score of annotator x over topic j asRSx,j,whereRSx,j = y∈W ∧TPy,j=0 agreement(x,y)”, & Page 5 – Section 3.4, “We start processing with top-1 instance in listW,weinferthelabelofthisinstanceusingtheinitialreliabilityscores, update the reliability scores for topics comprised in the top-1 instance according to its inferred label, then move to infer the label of top-2 instance employing the up dated reliability scores, reupdate again the reliability socres accordingly and so on”, thus updating a fidelity metric associated with each of […human evaluators…] is disclosed, because Jamaludeen teaches assigning a reliability score to each annotator and incrementally updating that reliability score according to the annotator’s vote relative to the inferred label. Jamaludeen’s annotator corresponds to a human evaluator, and the annotator’s reliability score corresponds to a fidelity metric associated with the human evaluator)
[…] and the updated fidelity metric […] (Jamaludeen, Page 4 – Section 3.4, “The votes are weighted with the annotators’ topic-based reliability scores”, & Page 5 – Section 3.4, “Each vote is weighted with the sum of the annotator tweet-related topic reliability scores. The weights are aggregated for annotators who provided identical votes by summing them up”, thus […] and the updated fidelity metric […] is disclosed, because Jamaludeen teaches weighting each annotator’s vote using the annotator’s topic-based reliability score and then aggregating the weighted votes. Jamaludeen’s topic-based reliability score corresponds to the updated fidelity metric, and using that score to weight each annotator’s vote corresponds to determining the result according to the updated fidelity metric)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Johnson with Jamaludeen by incorporating Jamaludeen’s annotator reliability scoring and weighted voting technique into Johnson’s Q&A user feedback system. Johnson teaches receiving feedback ratings from human operators or analysts and processing those ratings to determine confidence measures for generated candidate answers and refine the training of the Q&A system. Jamaludeen teaches assigning and updating reliability scores for annotators and weighting their evaluations according to those reliability scores. Therefore, a POSITA would have been motivated to use Jamaludeen’s reliability scores to weight the feedback ratings received from Johnson’s human operators so that feedback from more reliable evaluators would have a greater effect on the resulting answer ranking, while feedback from less reliable evaluators would have a reduced effect, thereby improving the accuracy and robustness of Johnson’s feedback processing, reducing the influence of unreliable human evaluations, and improving the reliability of the resulting answer evaluation and Q&A system performance (Jamaludeen, Page 11 – Section 5.6, “As a result, the more anno tations the annotator delivers, the model’s capacity of estimating the annotator’s re liability scores improves, thus the labels inference enhances. The model was also robust across different percentages of unreliable annotators and performed better than the Kappa Weighted Voting approach”)
Johnson combined with Jamaludeen does not explicitly teach updating a fidelity attribute associated with […the machine expert based on the cumulative ranking…], wherein the fidelity attribute is indicative of an ability of […the machine expert…] in answering questions in […the subject matter…], and […] and the updated fidelity attribute of […].
However, Wang teaches:
updating a fidelity attribute associated with […the machine expert based on the cumulative ranking…], wherein the fidelity attribute is indicative of an ability of […the machine expert…] in answering questions in […the subject matter…] (Wang, Par. [0101], “Assistant module 222 may determine a score of the experience, and feed the determined score back into ranking. For instance, assistant module 222 may modify the agent-quality score of the agent that fulfilled the request based on the user’s feedback about the fulfillment”, & Par. [0114], “the agent directory may collect user reviews and ratings. The collected user reviews and ratings may be used to modify the agent quality scores. As one example, when an agent receives positive reviews and/or ratings, agent accuracy module 331 may increase the agent’s agent quality score”, & Par. [0057], “local assistant module 122A may select a 3P agent based on rankings. For instance, local assistant module 122A may select a 3P agent with the highest ranking to satisfy the utterance”, thus updating a fidelity attribute associated with […the machine expert based on the cumulative ranking…], wherein the fidelity attribute is indicative of an ability of […the machine expert…] in answering questions in […the subject matter…] is disclosed, because Wang teaches determining an experience score from user feedback, feeding that score back into the ranking, and modifying the agent quality score of the particular agent that fulfilled the request. Wang further teaches increasing or otherwise modifying an agent’s quality score based on user reviews and ratings and selecting an agent based on its ranking for satisfying an utterance. Wang’s agent corresponds to a machine expert, the agent-quality score corresponds to a fidelity attribute, and the ranking used to select an agent for satisfying an utterance corresponds to an indication of the agent’s ability with respect to the subject matter)
[…] and the updated fidelity attribute of […] (Wang, Par. [0101], “assistant module 222 may modify the agent-quality score of the agent that fulfilled the request based on the user’s feedback about the fulfillment”, & Par. [0114], “The collected user reviews and ratings may be used to modify the agent quality scores”, thus […] and the updated fidelity attribute of […] is disclosed, because Wang teaches modifying an existing agent quality score based on user feedback and further modifying agent quality scores based on collected user reviews and ratings. Wang’s modified agent quality score corresponds to the updated fidelity attribute)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Johnson and Jamaludeen with Wang by incorporating Wang’s agent quality scoring and agent selection technique into the combined Q&A user feedback system. Johnson and Jamaludeen teach processing human evaluations while accounting for evaluator reliability to provide a more reliable evaluation of generated answers. Wang teaches evaluating and ranking individual agents so that an appropriate agent may be selected for a particular task. Therefore, a POSITA would have been motivated to use the reliability weighted evaluation provided by Johnson and Jamaludeen to update the quality score of the agent that generated the evaluated answer and use that score in selecting an appropriate agent for subsequent questions, thereby improving user machine interaction, increasing the likelihood that an appropriate agent is selected for a particular task, and reducing unnecessary consumption of computing resources and power by less appropriate agents (Wang, Par. [0004], “a computational assistant may initially process the utterance and select an appropriate agent to respond to the utterance. That is, the most appropriate agent may be selected for any given task received in an utterance. This provides an adaptive interface which improves user-machine interaction. Having a single assistant initially process the utterance before an appropriate agent responds to the utterance prevents other (less appropriate) agents from wasting computing resources or consuming power, processing the utterance. Accordingly, the techniques of this disclosure may enable a system to use less power and system resources than other systems that enable multiple agents to simultaneously process received utterances”)
Regarding Claim 9, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 8 as cited above and Johnson further teaches:
wherein the Q&A system includes a plurality of machine experts for automatically generating answers to questions […] at least some of the plurality of machine experts […] associated with each of the plurality of machine experts (Johnson, Par. [0071], “The QA system may comprise multiple engines or modules comprising logic for performing various operations for processing an input question in a natural language, searching a corpus of information for generating candidate answers to the input question, ranking or scoring the candidate answers, and performing a final merging of the scored or ranked candidate answers to generate a single ultimate answer to the input question,” thus wherein the Q&A system includes a plurality of machine experts for automatically generating answers to questions […] at least some of the plurality of machine experts […] associated with each of the plurality of machine experts is disclosed, because Johnson teaches a Q&A system having multiple engines or modules that process input questions, generate candidate answers, rank or score those candidate answers, and merge the ranked or scored answers to generate an ultimate answer. Johnson’s multiple engines or modules correspond to the plurality of machine experts, and the candidate answer generation performed by the engines or modules corresponds to automatically generating answers to questions)
Wang further teaches:
[…], wherein, for each question asked, […] are selected for providing an answer to the question and the selection is based, at least partially, on the fidelity attribute […] (Wang, Par. [0053], “local assistant module 122 may select an agent from a plurality of agents to satisfy the utterance”, & Par. [0057], “local assistant module 122A may select a 3P agent with the highest ranking to satisfy the utterance. In some examples, such as where there is a tie in the rankings and/or if the ranking of the 3P agent with the highest ranking is less than a ranking threshold, local assistant module 122A may solicit user input to select a 3P language agent to satisfy the utterance”, & Par. [0094], “Agent selection module 227 may rank the identified agent documents (e.g., based on a capability level to satisfy the utterance). For instance, agent selection module 227 may multiply a text-match score with an agent-quality score,” thus […], wherein, for each question asked, […] are selected for providing an answer to the question and the selection is based, at least partially, on the fidelity attribute […] is disclosed, because Wang teaches selecting an agent from a plurality of agents to satisfy an utterance, selecting a highest ranked agent to satisfy the utterance, and ranking candidate agents based on a capability level to satisfy the utterance using an agent-quality score. Wang’s selection of an agent to satisfy the utterance corresponds to selecting for providing an answer to the question, and Wang’s use of the agent quality score in ranking the agents corresponds to basing the selection, at least partially, on the fidelity attribute)
Regarding Claim 10, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 8 as cited above and Johnson further teaches:
identifying a ranking for the answer provided by the human evaluator from the evaluation (Johnson, Par. [0100], “The user may specify, for example, that the candidate answer "George Washington" is the most useful answer, or the answer having the highest quality, relative to the other candidate answers, and that the answers "General Washington" and "Washington" have relatively lower and lowest user ratings”, thus identifying a ranking for the answer provided by the human evaluator from the evaluation is disclosed, because Johnson teaches receiving a user evaluation that identifies one candidate answer as having the highest quality relative to other candidate answers and assigns relatively lower ratings to the remaining answers. Johnson’s relative user ratings correspond to the ranking for the answer identified from the human evaluator’s evaluation)
Jamaludeen further teaches:
determining a number of other rankings from remainder of […the one or more human evaluators…] that are consistent with […the ranking…] (Jamaludeen, Page 4 – Section 3.2, “For tweet y and class label c, let votes(y,c) be the number of annotators who as signed c to y. We assign each tweet to the class according to the majority voting, i.e. mvlabel(y) = argmaxc∈Cvotes(y,c).We use this number also to assign a rank to y” & “The rank reflects the agreement of annotators concerning the selected class label according to the majority voting labelling. We consider consensus as indicator of how much the class label of the instance can be trusted, and process high-ranked instances before low-ranked instances when computing annotator reliability”, thus determining a number of other rankings from remainder of […the one or more human evaluators…] that are consistent with […the ranking…] is disclosed, because Jamaludeen teaches counting the number of annotators that assign a particular class label and using that number to determine a majority label and corresponding rank reflecting agreement among the annotators. Jamaludeen’s annotators correspond to […the one or more human evaluators…], the annotator votes correspond to the rankings, and the number of annotators assigning the same class label corresponds to the number of other rankings that are consistent with […the ranking…])
determining a parameter based on the number of rankings from others (Jamaludeen, Page 4 – Section 3.2, “We assign each tweet to the class according to the majority voting, i.e. mvlabel(y) = argmaxc∈Cvotes(y,c).We use this number also to assign a rank to y”, & Page 4 – Section 3.3, “To distinguish reliable annotators from unreliable ones, we introduce the concept of reliability score: for each tweet y ∈W annotated by x, we set agreement(x,y)= 0, if vote(x,y)=inferredlabel(y) 1, if vote(x,y)=inferredlabel(y)”, thus determining a parameter based on the number of rankings from others is disclosed, because Jamaludeen teaches determining an inferred label from the number of annotator votes and then determining an agreement value according to whether an annotator’s vote is consistent with the inferred label. Jamaludeen’s agreement value corresponds to the parameter, and because the inferred label is determined from the number of votes from the other annotators, the agreement value is based on the number of rankings from others)
updating an existing fidelity metric associated with […the human evaluator…] based on the parameter to generate the updated fidelity metric for […the human evaluator…] (Jamaludeen, Page 4 – Section 3.4, “The tweets y ∈ W are processed incrementally and the reliability scores are updated simultaneously”, & Pge 5 – Section 3.4, “For each annotator who gave a vote identical to the inferred label, increment the tweet-related topic reliability scores by 1 as follows: RSx,j,t ←RSx,j,t−1+1”, thus updating an existing fidelity metric associated with […the human evaluator…] based on the parameter to generate the updated fidelity metric for […the human evaluator…] is disclosed, because Jamaludeen teaches incrementally updating annotator reliability scores and increasing the reliability score of an annotator when the annotator’s vote is identical to the inferred label. Jamaludeen’s annotator corresponds to […the human evaluator…], the existing reliability score corresponds to the existing fidelity metric, the determination that the annotator’s vote matches the inferred label corresponds to the parameter, and the incremented reliability score corresponds to the updated fidelity metric for […the human evaluator…])
Regarding Claim 11, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 8 as cited above and Johnson further teaches:
retrieving an existing ranking for the answer for the question (Johnson, Par. [0076], “The training run results query engine 510 provides logic for allowing users to use the QA system 500 to answer questions about the results generated by the QA system 500 of a previous run, or execution, of the QA system on a previously submitted question”, & Par. [0077], “results from inputting a training question to the QA system 500 are generated by the QA system 500 processing a training set of data to generate a set of candidate answers with corresponding confidence measures”, thus retrieving an existing ranking for the answer for the question is disclosed, because Johnson teaches accessing results from a previous execution of the QA system for a previously submitted question, wherein the results include candidate answers having corresponding confidence measures. Johnson’s previously generated confidence measure associated with a candidate answer corresponds to the existing ranking for the answer for the question)
determining the cumulative ranking of the answer based on the existing ranking […] (Johnson, Par. [0098], “the user feedback ratings of the candidate answers may further be used to modify the logic or configuration parameters so as to increase the confidence measures of the answers the users found to be most useful or having the highest quality while reducing those that have a relatively lower user feedback rating”), thus determining the cumulative ranking of the answer based on the existing ranking […] is disclosed, because Johnson teaches modifying an existing confidence measure associated with an answer based on user feedback ratings. Johnson’s existing confidence measure corresponds to the existing ranking, and the modified confidence measure corresponds to the cumulative ranking)
Jamaludeen further teaches:
accessing the updated fidelity metric for each of […the one or more human evaluators…] (Jamaludeen, Page 5 – Section 3.4, “we infer the label of this instance using the initial reliability scores, update the reliability scores for topics comprised in the top-1 instance according to its inferred label, then move to infer the label of top-2 instance employing the up dated reliability scores”, thus accessing the updated fidelity metric for each of […the one or more human evaluators…] is disclosed, because Jamaludeen teaches using previously updated annotator reliability scores when processing a subsequent instance. Jamaludeen’s annotators correspond to […the one or more human evaluators…], and the updated annotator reliability scores correspond to the updated fidelity metrics associated with those evaluators)
identifying a ranking from the evaluation from each of [… the one or more human evaluators...] (Jamaludeen, Page 5 – Section 3.4, “Each vote is weighted with the sum of the annotator tweet related topic reliability scores”, & “The weights are aggregated for annotators who provided identical votes by summing them up”, thus identifying a ranking from the evaluation from each of […the one or more human evaluators…] is disclosed, because Jamaludeen teaches identifying the individual vote provided by each annotator before weighting and aggregating the votes. Jamaludeen’s annotators correspond to […the one or more human evaluators…], and each annotator’s vote corresponds to a ranking identified from that evaluator’s evaluation)
weighing the ranking […of each of the one or more human evaluators…] based on the updated fidelity metric thereof to generate a weighted ranking for […the human evaluator…] (Jamaludeen, Page 5 – Section 3.4, “Each vote is weighted with the sum of the annotator tweet related topic reliability scores”, thus weighing the ranking […of each of the one or more human evaluators…] based on the updated fidelity metric thereof to generate a weighted ranking for […the human evaluator…] is disclosed, because Jamaludeen teaches weighting each annotator’s vote using the annotator’s reliability scores. Jamaludeen’s annotator corresponds to […the one or more human evaluators…], the annotator’s vote corresponds to the ranking, the annotator’s reliability score corresponds to the updated fidelity metric, and the resulting weighted vote corresponds to the weighted ranking for […the human evaluator…])
obtaining an integrated ranking for the answer based on the weighted ranking of […each of the one or more human evaluators…] (Jamaludeen, Page 5 – Section 3.4, “The weights are aggregated for annotators who provided identical votes by summing them up”, & “We select the class label that collected the highest weight as the label for the tweet”, thus obtaining an integrated ranking for the answer based on the weighted ranking of […each of the one or more human evaluators…] is disclosed, because Jamaludeen teaches aggregating the weighted votes of multiple annotators and determining the resulting outcome based on the accumulated weights. Jamaludeen’s weighted annotator votes correspond to the weighted rankings of […each of the one or more human evaluators…], and the aggregated weight used to determine the resulting label corresponds to the integrated ranking for the answer)
[…] and the integrated ranking for […the answer…] (Jamaludeen, Page 5 - Section 3.4, “The weights are aggregated for annotators who provided identical votes by summing them up”, & “We select the class label that collected the highest weight as the label for the tweet”), thus […] and the integrated ranking for […the answer…] is disclosed, because Jamaludeen teaches aggregating the weighted votes of multiple annotators and determining the resulting outcome based on the accumulated weights. Jamaludeen’s aggregated weighted votes correspond to the integrated ranking for […the answer…])
Regarding Claim 12, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 11 as cited above and Wang further teaches:
retrieving an existing fidelity attribute associated with […the machine expert…] (Wang, Par. [0047], “agent indices 124 may store information related to the use and/or the performance of the available agents. For instance, agent indices 124 may include an agent-quality score for each available agent”, thus retrieving an existing fidelity attribute associated with […the machine expert…] is disclosed, because Wang teaches storing an agent-quality score associated with each available agent. Wang’s agent corresponds to […the machine expert…], and the stored agent-quality score corresponds to the existing fidelity attribute associated with […the machine expert…])
modifying the existing fidelity attribute based on […the cumulative ranking of the answer…] (Wang, Par. [0101], “assistant module 222 may modify the agent-quality score of the agent that fulfilled the request based on the user's feedback about the fulfillment”, & Par. [0114], “the agent directory may collect user reviews and ratings. The collected user reviews and ratings may be used to modify the agent quality scores”, thus modifying the existing fidelity attribute based on […the cumulative ranking of the answer…] is disclosed, because Wang teaches modifying an existing agent-quality score based on feedback and ratings concerning the performance of the agent. Wang’s existing agent-quality score corresponds to the existing fidelity attribute)
generating an updated fidelity attribute based on the modified fidelity attribute for […the machine expert…] (Wang, Par. [0114], “when an agent receives positive reviews and/or ratings, agent accuracy module 331 may increase the agent's agent quality score in agent index 224 or agent index 324. As another example, when an agent receives negative reviews and/or ratings, agent accuracy module 331 may decrease the agent's agent quality score in agent index 224 or agent index 324”, thus generating an updated fidelity attribute based on the modified fidelity attribute for […the machine expert…] is disclosed, because Wang teaches modifying an existing agent quality score by increasing or decreasing the score based on received reviews and ratings, thereby producing an updated agent quality score. Wang’s updated agent quality score corresponds to the updated fidelity attribute for […the machine expert…])
Regarding Claim 13, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 8 as cited above and Johnson further teaches:
wherein the feedback further includes at least one of an alternative answer in place of the answer provided by one of the one or more human evaluators, an alternative reference that supports the alternative answer, and an alternative source to access the alternative reference (Johnson, Par. [0093], “The user feedback rating engine 720 provides logic for presenting a listing of candidate answers generated by the QA system 710 and provided in the results 715, as well as their confidence measures, answer source information, e.g., corpus, document, etc., and the like, via a user interface 725”, & Par. [0100], “the user is presented with a listing of the candidate answers and options or fields through which a user may specify a rating of each of the candidate answers, via the user feedback rating engine 720”, thus wherein the feedback further includes at least one of an alternative answer in place of the answer provided by one of the one or more human evaluators, an alternative reference that supports the alternative answer, and an alternative source to access the alternative reference is disclosed, because Johnson teaches presenting multiple candidate answers and corresponding answer source information to a user and receiving user feedback ratings concerning those candidate answers. Johnson’s candidate answers other than a particular answer correspond to an alternative answer in place of the answer, and Johnson’s corpus or document information associated with the candidate answers corresponds to an alternative reference or alternative source supporting the alternative answer)
Regarding Claim 14, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 13 as cited above and Johnson further teaches:
extracting, from the feedback, information related to the alternative reference and the alternative source (Johnson, Par. [0093], “The user feedback rating engine 720 further provides the user interface 725 with elements and/or fields through which user input is received for rating the quality of the candidate answers and/or answer source information presented in the results 715”, & Par. [0102], “the QA system 710 may identify the sources upon which the QA system 710 relies for each of the candidate answers in the results 715”, thus extracting, from the feedback, information related to the alternative reference and the alternative source is disclosed, because Johnson teaches receiving user feedback directed to candidate answers and associated answer source information and identifying the sources relied upon for the candidate answers. Johnson’s answer source information corresponds to information related to the alternative reference and the alternative source, and identifying that source information from the feedback corresponds to extracting the information related to the alternative reference and alternative source)
modifying an archive storing references from different sources based on the alternative reference and the alternative source (Johnson, Par. [0074], “the document/passage retrieval module 503 applies the queries to a corpus of information 502, such as a database of structure and/or unstructured document data, and performs document filtering and passage post-filtering in a document containing the content, e.g., keywords matching criteria of one or more of the queries, so as to generate candidate answers”, & Par. [0102], “the user feedback, obtained via the user feedback rating engine 720, may be accumulated for an answer source and may be used to modify the confidence measures associated with the answer source”, thus modifying an archive storing references from different sources based on the alternative reference and the alternative source is disclosed, because Johnson teaches using a corpus of information comprising structured or unstructured document data as the source from which candidate answers are generated, and further teaches modifying information associated with an answer source based on accumulated user feedback. Johnson’s corpus of information corresponds to the archive storing references from different sources, the documents in the corpus correspond to the references, and modifying the confidence measures associated with an answer source based on feedback corresponds to modifying information in the archive based on the alternative reference and the alternative source)
Regarding Claim 15, Johnson teaches:
receiving, from one or more human evaluators, evaluation directed to an answer automatically generated by a machine expert in a question & answer (Q&A) system in response to a question related to a subject matter based on a reference from a source (Johnson, Par. [0093], “The user feedback rating engine 720 provides logic for presenting a listing of candidate answers generated by the QA system 710 and provided in the results 715, as well as their confidence measures, answer source information, e.g., corpus, document, etc., and the like, via a user interface 725. The user feedback rating engine 720 further provides the user interface 725 with elements and/or fields through which user input is received for rating the quality of the candidate answers and/or answer source information presented in the results 715”, & Par. [0071], “The QA system may comprise multiple engines or modules comprising logic for performing various operations for processing an input question in a natural language, searching a corpus of information for generating candidate answers to the input question, ranking or scoring the candidate answers, and performing a final merging of the scored or ranked candidate answers to generate a single ultimate answer to the input question. Thus, the QA system may comprise engines/modules for performing question analysis, content analysis of documents in a corpus of information, primary search, candidate answer generation, candidate answer scoring/ranking, and final merging of candidate answers”, & Par. [0074], “For a natural language question raised by a user, question parsing and focus detecting are performed in the question processing module 501, which generates queries for the question”, & Par. [0074], “Afterwards, the answer processing module 505 performs candidate identification and answer ranking on the candidate answers generated by the document/passage retrieval module 503, and finally formulates an answer to the raised natural language question, so as to output a brief answer to the user in natural language”, thus Johnson teaches a QA system having multiple engines/modules configured to process an input question, search a corpus of information, generate candidate answers, and formulate an answer to the question, and further teaches receiving user input that rates the quality of the generated candidate answers and associated answer-source information. Johnson’s candidate-answer-generation engines/modules correspond to a machine expert because the engines/modules automatically generate candidate answers in response to an input question. Johnson’s candidate answers correspond to an answer, the user providing a rating corresponds to a human evaluator, the user rating corresponds to evaluation, and Johnson’s corpus/document and answer source information correspond to a reference from a source)
[…] the one or more human evaluators based on the evaluation (Johnson, Par. [0031], “human operators, analysts, or the like, are able to provide feedback ratings of candidate answers to posed questions during the training process and these feedback ratings are automatically processed in addition to the processing of the generated candidate answers to determine confidence measures for the generated candidate answers and refine a training model for performing machine learning during the training of the QA systems”, thus […] the one or more human evaluators based on the evaluation is disclosed, because Johnson teaches that human operators or analysts provide feedback ratings of candidate answers and that those feedback ratings are automatically processed to determine confidence measures and refine the training model. Johnson’s human operators or analysts correspond to the one or more human evaluators, and the feedback ratings correspond to the evaluation)
determining a cumulative ranking of the answer with respect to the question according to the evaluation […] of each of the one or more human evaluators (Johnson, Par. [0098], “the user feedback ratings of the candidate answers may further be used to modify the logic or configuration parameters so as to increase the confidence measures of the answers the users found to be most useful or having the highest quality while reducing those that have a relatively lower user feedback rating”, thus determining a cumulative ranking of the answer with respect to the question according to the evaluation […] of each of the one or more human evaluators is disclosed, because Johnson teaches using user feedback ratings of candidate answers to increase or decrease the confidence measures associated with the answers according to the users’ assessment of their usefulness or quality. Johnson’s user feedback ratings correspond to the evaluation, and the confidence measures associated with the candidate answers correspond to the ranking of the answer)
generating feedback based on the answer, the question, the cumulative ranking of the answer with respect to the question, […] the machine expert (Johnson, Par. [0073], “The illustrative embodiments provide mechanisms for automatically identifying areas of improvement, provide mechanisms for generating recommendations for such improvement, and provide mechanisms for actually improving the operation of a QA system with regard to the accuracy or confidence measures of candidate answers generated by the QA system. As mentioned previously, the illustrative embodiments provide mechanisms for answering questions posed regarding the results generated by the training of the QA system, mechanisms for allowing users to input ratings of the candidate answers which may then be processed by the QA system to determine improvements to the model or logic used by the QA system”, & Par. [0098], “the user feedback ratings of the candidate answers may further be used to modify the logic or configuration parameters so as to increase the confidence measures of the answers the users found to be most useful or having the highest quality while reducing those that have a relatively lower user feedback rating”, & Par. [0099], “the user poses the training question "Who was the man that crossed the Potomac?" There may be multiple candidate answers that may be returned by the QA system 710 in the results 715 with separate confidence measures for each of these candidate answers”, thus generating feedback based on the answer, the question, the cumulative ranking of the answer with respect to the question, […] the machine expert is disclosed, because Johnson teaches generating recommendations for improving the QA system based on candidate answers, the corresponding questions, user ratings of the candidate answers, and confidence measures associated with those answers. Johnson’s generated recommendations correspond to feedback, the candidate answers correspond to the answer, the posed training question corresponds to the question, and the confidence measures adjusted according to the user feedback ratings correspond to the cumulative ranking of the answer with respect to the question)
sending the feedback to the Q&A system for adapting the Q&A system (Johnson, Par. [0073], “mechanisms for allowing users to input ratings of the candidate answers which may then be processed by the QA system to determine improvements to the model or logic used by the QA system”, & Par. [0096], “the user rating feedback information may be used to refine the training performed by the training engine 740 based on the identified differences between the results 715 and the correct answers specified in the ground truth data storage 730. The modifications to the QA system 710 based on user rating feedback may be implemented, for example, in various ones of the modules/engines of the QA system 710”, & Par. [0101], “Thus, the user feedback may be used to adjust the training of the QA system 710”, thus sending the feedback to the Q&A system for adapting the Q&A system is disclosed, because Johnson teaches providing user ratings of candidate answers to the QA system, processing those ratings to determine improvements to the QA system’s model or logic, and using the feedback to refine training and modify modules or engines of the QA system. Johnson’s user rating feedback corresponds to the feedback, and refining the training and modifying the model, logic, or modules of the QA system corresponds to adapting the Q&A system)
Johnson does not explicitly teach updating a fidelity metric associated with each of […human evaluators…], […] and the updated fidelity metric […], updating a fidelity metric associated with each of […human evaluators…], updating a fidelity attribute associated with […the machine expert based on the cumulative ranking…], wherein the fidelity attribute is indicative of an ability of […the machine expert…] in answering questions in […the subject matter…], and […] and the updated fidelity attribute of […].
However, Jamaludeen teaches:
updating a fidelity metric associated with each of […human evaluators…] (Jamaludeen, Page 4 – Section 3.3, “To distinguish reliable annotators from unreliable ones, we introduce the concept of reliability score: for each tweet y ∈W annotated by x, we set agreement(x,y)= 0, if vote(x,y) = inferred label(y) 1, if vote(x,y) = inferred label(y) (1) Then,we define the reliability score of annotator x over topic j asRSx,j,whereRSx,j = y∈W ∧TPy,j=0 agreement(x,y)”, & Page 5 – Section 3.4, “We start processing with top-1 instance in listW,weinferthelabelofthisinstanceusingtheinitialreliabilityscores, update the reliability scores for topics comprised in the top-1 instance according to its inferred label, then move to infer the label of top-2 instance employing the up dated reliability scores, reupdate again the reliability socres accordingly and so on”, thus updating a fidelity metric associated with each of […human evaluators…] is disclosed, because Jamaludeen teaches assigning a reliability score to each annotator and incrementally updating that reliability score according to the annotator’s vote relative to the inferred label. Jamaludeen’s annotator corresponds to a human evaluator, and the annotator’s reliability score corresponds to a fidelity metric associated with the human evaluator)
[…] and the updated fidelity metric […] (Jamaludeen, Page 4 – Section 3.4, “The votes are weighted with the annotators’ topic-based reliability scores”, & Page 5 – Section 3.4, “Each vote is weighted with the sum of the annotator tweet-related topic reliability scores. The weights are aggregated for annotators who provided identical votes by summing them up”, thus […] and the updated fidelity metric […] is disclosed, because Jamaludeen teaches weighting each annotator’s vote using the annotator’s topic-based reliability score and then aggregating the weighted votes. Jamaludeen’s topic-based reliability score corresponds to the updated fidelity metric, and using that score to weight each annotator’s vote corresponds to determining the result according to the updated fidelity metric)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to combine Johnson with Jamaludeen by incorporating Jamaludeen’s annotator reliability scoring and weighted voting technique into Johnson’s Q&A user feedback system. Johnson teaches receiving feedback ratings from human operators or analysts and processing those ratings to determine confidence measures for generated candidate answers and refine the training of the Q&A system. Jamaludeen teaches assigning and updating reliability scores for annotators and weighting their evaluations according to those reliability scores. Therefore, a POSITA would have been motivated to use Jamaludeen’s reliability scores to weight the feedback ratings received from Johnson’s human operators so that feedback from more reliable evaluators would have a greater effect on the resulting answer ranking, while feedback from less reliable evaluators would have a reduced effect, thereby improving the accuracy and robustness of Johnson’s feedback processing, reducing the influence of unreliable human evaluations, and improving the reliability of the resulting answer evaluation and Q&A system performance (Jamaludeen, Page 11 – Section 5.6, “As a result, the more anno tations the annotator delivers, the model’s capacity of estimating the annotator’s re liability scores improves, thus the labels inference enhances. The model was also robust across different percentages of unreliable annotators and performed better than the Kappa Weighted Voting approach”)
Johnson combined with Jamaludeen does not explicitly teach updating a fidelity attribute associated with […the machine expert based on the cumulative ranking…], wherein the fidelity attribute is indicative of an ability of […the machine expert…] in answering questions in […the subject matter…], and […] and the updated fidelity attribute of […].
However, Wang teaches:
updating a fidelity attribute associated with […the machine expert based on the cumulative ranking…], wherein the fidelity attribute is indicative of an ability of […the machine expert…] in answering questions in […the subject matter…] (Wang, Par. [0101], “Assistant module 222 may determine a score of the experience, and feed the determined score back into ranking. For instance, assistant module 222 may modify the agent-quality score of the agent that fulfilled the request based on the user’s feedback about the fulfillment”, & Par. [0114], “the agent directory may collect user reviews and ratings. The collected user reviews and ratings may be used to modify the agent quality scores. As one example, when an agent receives positive reviews and/or ratings, agent accuracy module 331 may increase the agent’s agent quality score”, & Par. [0057], “local assistant module 122A may select a 3P agent based on rankings. For instance, local assistant module 122A may select a 3P agent with the highest ranking to satisfy the utterance”, thus updating a fidelity attribute associated with […the machine expert based on the cumulative ranking…], wherein the fidelity attribute is indicative of an ability of […the machine expert…] in answering questions in […the subject matter…] is disclosed, because Wang teaches determining an experience score from user feedback, feeding that score back into the ranking, and modifying the agent quality score of the particular agent that fulfilled the request. Wang further teaches increasing or otherwise modifying an agent’s quality score based on user reviews and ratings and selecting an agent based on its ranking for satisfying an utterance. Wang’s agent corresponds to a machine expert, the agent-quality score corresponds to a fidelity attribute, and the ranking used to select an agent for satisfying an utterance corresponds to an indication of the agent’s ability with respect to the subject matter)
[…] and the updated fidelity attribute of […] (Wang, Par. [0101], “assistant module 222 may modify the agent-quality score of the agent that fulfilled the request based on the user’s feedback about the fulfillment”, & Par. [0114], “The collected user reviews and ratings may be used to modify the agent quality scores”, thus […] and the updated fidelity attribute of […] is disclosed, because Wang teaches modifying an existing agent quality score based on user feedback and further modifying agent quality scores based on collected user reviews and ratings. Wang’s modified agent quality score corresponds to the updated fidelity attribute)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to further combine Johnson and Jamaludeen with Wang by incorporating Wang’s agent quality scoring and agent selection technique into the combined Q&A user feedback system. Johnson and Jamaludeen teach processing human evaluations while accounting for evaluator reliability to provide a more reliable evaluation of generated answers. Wang teaches evaluating and ranking individual agents so that an appropriate agent may be selected for a particular task. Therefore, a POSITA would have been motivated to use the reliability weighted evaluation provided by Johnson and Jamaludeen to update the quality score of the agent that generated the evaluated answer and use that score in selecting an appropriate agent for subsequent questions, thereby improving user machine interaction, increasing the likelihood that an appropriate agent is selected for a particular task, and reducing unnecessary consumption of computing resources and power by less appropriate agents (Wang, Par. [0004], “a computational assistant may initially process the utterance and select an appropriate agent to respond to the utterance. That is, the most appropriate agent may be selected for any given task received in an utterance. This provides an adaptive interface which improves user-machine interaction. Having a single assistant initially process the utterance before an appropriate agent responds to the utterance prevents other (less appropriate) agents from wasting computing resources or consuming power, processing the utterance. Accordingly, the techniques of this disclosure may enable a system to use less power and system resources than other systems that enable multiple agents to simultaneously process received utterances”)
Regarding Claim 16, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 15 as cited above and Johnson further teaches:
wherein the Q&A system includes a plurality of machine experts for automatically generating answers to questions […] at least some of the plurality of machine experts […] associated with each of the plurality of machine experts (Johnson, Par. [0071], “The QA system may comprise multiple engines or modules comprising logic for performing various operations for processing an input question in a natural language, searching a corpus of information for generating candidate answers to the input question, ranking or scoring the candidate answers, and performing a final merging of the scored or ranked candidate answers to generate a single ultimate answer to the input question,” thus wherein the Q&A system includes a plurality of machine experts for automatically generating answers to questions […] at least some of the plurality of machine experts […] associated with each of the plurality of machine experts is disclosed, because Johnson teaches a Q&A system having multiple engines or modules that process input questions, generate candidate answers, rank or score those candidate answers, and merge the ranked or scored answers to generate an ultimate answer. Johnson’s multiple engines or modules correspond to the plurality of machine experts, and the candidate answer generation performed by the engines or modules corresponds to automatically generating answers to questions)
Wang further teaches:
[…], wherein, for each question asked, […] are selected for providing an answer to the question and the selection is based, at least partially, on the fidelity attribute […] (Wang, Par. [0053], “local assistant module 122 may select an agent from a plurality of agents to satisfy the utterance”, & Par. [0057], “local assistant module 122A may select a 3P agent with the highest ranking to satisfy the utterance. In some examples, such as where there is a tie in the rankings and/or if the ranking of the 3P agent with the highest ranking is less than a ranking threshold, local assistant module 122A may solicit user input to select a 3P language agent to satisfy the utterance”, & Par. [0094], “Agent selection module 227 may rank the identified agent documents (e.g., based on a capability level to satisfy the utterance). For instance, agent selection module 227 may multiply a text-match score with an agent-quality score,” thus […], wherein, for each question asked, […] are selected for providing an answer to the question and the selection is based, at least partially, on the fidelity attribute […] is disclosed, because Wang teaches selecting an agent from a plurality of agents to satisfy an utterance, selecting a highest ranked agent to satisfy the utterance, and ranking candidate agents based on a capability level to satisfy the utterance using an agent-quality score. Wang’s selection of an agent to satisfy the utterance corresponds to selecting for providing an answer to the question, and Wang’s use of the agent quality score in ranking the agents corresponds to basing the selection, at least partially, on the fidelity attribute)
Regarding Claim 17, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 15 as cited above and Johnson further teaches:
identifying a ranking for the answer provided by the human evaluator from the evaluation (Johnson, Par. [0100], “The user may specify, for example, that the candidate answer "George Washington" is the most useful answer, or the answer having the highest quality, relative to the other candidate answers, and that the answers "General Washington" and "Washington" have relatively lower and lowest user ratings”, thus identifying a ranking for the answer provided by the human evaluator from the evaluation is disclosed, because Johnson teaches receiving a user evaluation that identifies one candidate answer as having the highest quality relative to other candidate answers and assigns relatively lower ratings to the remaining answers. Johnson’s relative user ratings correspond to the ranking for the answer identified from the human evaluator’s evaluation)
Jamaludeen further teaches:
determining a number of other rankings from remainder of […the one or more human evaluators…] that are consistent with […the ranking…] (Jamaludeen, Page 4 – Section 3.2, “For tweet y and class label c, let votes(y,c) be the number of annotators who as signed c to y. We assign each tweet to the class according to the majority voting, i.e. mvlabel(y) = argmaxc∈Cvotes(y,c).We use this number also to assign a rank to y” & “The rank reflects the agreement of annotators concerning the selected class label according to the majority voting labelling. We consider consensus as indicator of how much the class label of the instance can be trusted, and process high-ranked instances before low-ranked instances when computing annotator reliability”, thus determining a number of other rankings from remainder of […the one or more human evaluators…] that are consistent with […the ranking…] is disclosed, because Jamaludeen teaches counting the number of annotators that assign a particular class label and using that number to determine a majority label and corresponding rank reflecting agreement among the annotators. Jamaludeen’s annotators correspond to […the one or more human evaluators…], the annotator votes correspond to the rankings, and the number of annotators assigning the same class label corresponds to the number of other rankings that are consistent with […the ranking…])
determining a parameter based on the number of rankings from others (Jamaludeen, Page 4 – Section 3.2, “We assign each tweet to the class according to the majority voting, i.e. mvlabel(y) = argmaxc∈Cvotes(y,c).We use this number also to assign a rank to y”, & Page 4 – Section 3.3, “To distinguish reliable annotators from unreliable ones, we introduce the concept of reliability score: for each tweet y ∈W annotated by x, we set agreement(x,y)= 0, if vote(x,y)=inferredlabel(y) 1, if vote(x,y)=inferredlabel(y)”, thus determining a parameter based on the number of rankings from others is disclosed, because Jamaludeen teaches determining an inferred label from the number of annotator votes and then determining an agreement value according to whether an annotator’s vote is consistent with the inferred label. Jamaludeen’s agreement value corresponds to the parameter, and because the inferred label is determined from the number of votes from the other annotators, the agreement value is based on the number of rankings from others)
updating an existing fidelity metric associated with […the human evaluator…] based on the parameter to generate the updated fidelity metric for […the human evaluator…] (Jamaludeen, Page 4 – Section 3.4, “The tweets y ∈ W are processed incrementally and the reliability scores are updated simultaneously”, & Pge 5 – Section 3.4, “For each annotator who gave a vote identical to the inferred label, increment the tweet-related topic reliability scores by 1 as follows: RSx,j,t ←RSx,j,t−1+1”, thus updating an existing fidelity metric associated with […the human evaluator…] based on the parameter to generate the updated fidelity metric for […the human evaluator…] is disclosed, because Jamaludeen teaches incrementally updating annotator reliability scores and increasing the reliability score of an annotator when the annotator’s vote is identical to the inferred label. Jamaludeen’s annotator corresponds to […the human evaluator…], the existing reliability score corresponds to the existing fidelity metric, the determination that the annotator’s vote matches the inferred label corresponds to the parameter, and the incremented reliability score corresponds to the updated fidelity metric for […the human evaluator…])
Regarding Claim 18, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 15 as cited above and Johnson further teaches:
retrieving an existing ranking for the answer for the question (Johnson, Par. [0076], “The training run results query engine 510 provides logic for allowing users to use the QA system 500 to answer questions about the results generated by the QA system 500 of a previous run, or execution, of the QA system on a previously submitted question”, & Par. [0077], “results from inputting a training question to the QA system 500 are generated by the QA system 500 processing a training set of data to generate a set of candidate answers with corresponding confidence measures”, thus retrieving an existing ranking for the answer for the question is disclosed, because Johnson teaches accessing results from a previous execution of the QA system for a previously submitted question, wherein the results include candidate answers having corresponding confidence measures. Johnson’s previously generated confidence measure associated with a candidate answer corresponds to the existing ranking for the answer for the question)
determining the cumulative ranking of the answer based on the existing ranking […] (Johnson, Par. [0098], “the user feedback ratings of the candidate answers may further be used to modify the logic or configuration parameters so as to increase the confidence measures of the answers the users found to be most useful or having the highest quality while reducing those that have a relatively lower user feedback rating”), thus determining the cumulative ranking of the answer based on the existing ranking […] is disclosed, because Johnson teaches modifying an existing confidence measure associated with an answer based on user feedback ratings. Johnson’s existing confidence measure corresponds to the existing ranking, and the modified confidence measure corresponds to the cumulative ranking)
Jamaludeen further teaches:
accessing the updated fidelity metric for each of […the one or more human evaluators…] (Jamaludeen, Page 5 – Section 3.4, “we infer the label of this instance using the initial reliability scores, update the reliability scores for topics comprised in the top-1 instance according to its inferred label, then move to infer the label of top-2 instance employing the up dated reliability scores”, thus accessing the updated fidelity metric for each of […the one or more human evaluators…] is disclosed, because Jamaludeen teaches using previously updated annotator reliability scores when processing a subsequent instance. Jamaludeen’s annotators correspond to […the one or more human evaluators…], and the updated annotator reliability scores correspond to the updated fidelity metrics associated with those evaluators)
identifying a ranking from the evaluation from each of [… the one or more human evaluators...] (Jamaludeen, Page 5 – Section 3.4, “Each vote is weighted with the sum of the annotator tweet related topic reliability scores”, & “The weights are aggregated for annotators who provided identical votes by summing them up”, thus identifying a ranking from the evaluation from each of […the one or more human evaluators…] is disclosed, because Jamaludeen teaches identifying the individual vote provided by each annotator before weighting and aggregating the votes. Jamaludeen’s annotators correspond to […the one or more human evaluators…], and each annotator’s vote corresponds to a ranking identified from that evaluator’s evaluation)
weighing the ranking […of each of the one or more human evaluators…] based on the updated fidelity metric thereof to generate a weighted ranking for […the human evaluator…] (Jamaludeen, Page 5 – Section 3.4, “Each vote is weighted with the sum of the annotator tweet related topic reliability scores”, thus weighing the ranking […of each of the one or more human evaluators…] based on the updated fidelity metric thereof to generate a weighted ranking for […the human evaluator…] is disclosed, because Jamaludeen teaches weighting each annotator’s vote using the annotator’s reliability scores. Jamaludeen’s annotator corresponds to […the one or more human evaluators…], the annotator’s vote corresponds to the ranking, the annotator’s reliability score corresponds to the updated fidelity metric, and the resulting weighted vote corresponds to the weighted ranking for […the human evaluator…])
obtaining an integrated ranking for the answer based on the weighted ranking of […each of the one or more human evaluators…] (Jamaludeen, Page 5 – Section 3.4, “The weights are aggregated for annotators who provided identical votes by summing them up”, & “We select the class label that collected the highest weight as the label for the tweet”, thus obtaining an integrated ranking for the answer based on the weighted ranking of […each of the one or more human evaluators…] is disclosed, because Jamaludeen teaches aggregating the weighted votes of multiple annotators and determining the resulting outcome based on the accumulated weights. Jamaludeen’s weighted annotator votes correspond to the weighted rankings of […each of the one or more human evaluators…], and the aggregated weight used to determine the resulting label corresponds to the integrated ranking for the answer)
[…] and the integrated ranking for […the answer…] (Jamaludeen, Page 5 - Section 3.4, “The weights are aggregated for annotators who provided identical votes by summing them up”, & “We select the class label that collected the highest weight as the label for the tweet”), thus […] and the integrated ranking for […the answer…] is disclosed, because Jamaludeen teaches aggregating the weighted votes of multiple annotators and determining the resulting outcome based on the accumulated weights. Jamaludeen’s aggregated weighted votes correspond to the integrated ranking for […the answer…])
Regarding Claim 19, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 18 as cited above and Wang further teaches:
retrieving an existing fidelity attribute associated with […the machine expert…] (Wang, Par. [0047], “agent indices 124 may store information related to the use and/or the performance of the available agents. For instance, agent indices 124 may include an agent-quality score for each available agent”, thus retrieving an existing fidelity attribute associated with […the machine expert…] is disclosed, because Wang teaches storing an agent-quality score associated with each available agent. Wang’s agent corresponds to […the machine expert…], and the stored agent-quality score corresponds to the existing fidelity attribute associated with […the machine expert…])
modifying the existing fidelity attribute based on […the cumulative ranking of the answer…] (Wang, Par. [0101], “assistant module 222 may modify the agent-quality score of the agent that fulfilled the request based on the user's feedback about the fulfillment”, & Par. [0114], “the agent directory may collect user reviews and ratings. The collected user reviews and ratings may be used to modify the agent quality scores”, thus modifying the existing fidelity attribute based on […the cumulative ranking of the answer…] is disclosed, because Wang teaches modifying an existing agent-quality score based on feedback and ratings concerning the performance of the agent. Wang’s existing agent-quality score corresponds to the existing fidelity attribute)
generating an updated fidelity attribute based on the modified fidelity attribute for […the machine expert…] (Wang, Par. [0114], “when an agent receives positive reviews and/or ratings, agent accuracy module 331 may increase the agent's agent quality score in agent index 224 or agent index 324. As another example, when an agent receives negative reviews and/or ratings, agent accuracy module 331 may decrease the agent's agent quality score in agent index 224 or agent index 324”, thus generating an updated fidelity attribute based on the modified fidelity attribute for […the machine expert…] is disclosed, because Wang teaches modifying an existing agent quality score by increasing or decreasing the score based on received reviews and ratings, thereby producing an updated agent quality score. Wang’s updated agent quality score corresponds to the updated fidelity attribute for […the machine expert…])
Regarding Claim 20, Johnson and Jamaludeen combined with Wang teaches all the limitations of claim 15 as cited above and Johnson further teaches:
extracting, from the feedback, information related to at least one of an alternative answer in place of the answer provided by one of the one or more human evaluators, an alternative reference relied on to derive the alterative answer, and an alternative source to access the alternative reference (Johnson, Par. [0093], “The user feedback rating engine 720 further provides the user interface 725 with elements and/or fields through which user input is received for rating the quality of the candidate answers and/or answer source information presented in the results 715”, & Par. [0102], “the QA system 710 may identify the sources upon which the QA system 710 relies for each of the candidate answers in the results 715”, thus extracting, from the feedback, information related to at least one of an alternative answer in place of the answer provided by one of the one or more human evaluators, an alternative reference relied on to derive the alterative answer, and an alternative source to access the alternative reference is disclosed, because Johnson teaches receiving user feedback concerning candidate answers and associated answer source information and identifying the sources relied upon for those candidate answers. Johnson’s candidate answer information corresponds to information related to an alternative answer, and Johnson’s answer source information corresponds to information related to an alternative reference or alternative source)
modifying an archive storing references from different sources based on the alternative reference and the alternative source (Johnson, Par. [0074], “the document/passage retrieval module 503 applies the queries to a corpus of information 502, such as a database of structure and/or unstructured document data, and performs document filtering and passage post-filtering in a document containing the content, e.g., keywords matching criteria of one or more of the queries, so as to generate candidate answers”, & Par. [0102], “the user feedback, obtained via the user feedback rating engine 720, may be accumulated for an answer source and may be used to modify the confidence measures associated with the answer source”, thus modifying an archive storing references from different sources based on the alternative reference and the alternative source is disclosed, because Johnson teaches using a corpus of information comprising structured or unstructured document data as the source from which candidate answers are generated, and further teaches modifying information associated with an answer source based on accumulated user feedback. Johnson’s corpus of information corresponds to the archive storing references from different sources, the documents in the corpus correspond to the references, and modifying the confidence measures associated with an answer source based on feedback corresponds to modifying information in the archive based on the alternative reference and the alternative source)
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MAHLIET ADMASU whose telephone number is (571)272-0034. The examiner can normally be reached Mon-Fri, 8am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571)270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/M.T.A./Examiner, Art Unit 2123 /ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123