Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: "the data store, to perform operations" in claim 13.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. This limitation is structurally tied to the feature the claim recites of a processing system. The instant specification describes processing system as ([¶0096] "a processing system 1304 including one or more processors. The processor(s) include one or more central processing units (CPUs)…". Therefore, the above limitation is modified by sufficient structure, material, or acts for performing the claimed function.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 13-16 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 13 recites the limitation "the plural-objective model" in the phrase “using the plural-objective model to identify a set of". There is insufficient antecedent basis for this limitation in the claim.
Claim 13 further recites the limitation "the respective source items" in the phrase “that are relevant to the respective source items and which. There is insufficient antecedent basis for this limitation in the claim.
Claim 13 further recites the limitation "queries" in the phrase “based on queries that are relevant to. There is insufficient antecedent basis for this limitation in the claim. For purposes of examination, this expression has been interpreted as referring to the plurality of training-time queries and their corresponding source items used when the plural-objective model’s parameters were trained.
Claims 14-16 are rejected as dependent upon a rejected base claim.
Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Is the claim to a process, machine, manufacture or composition of matter?
Claims 1-12 are directed to a method (i.e., a process); claims 13-16 are directed
to an apparatus (i.e., a machine/apparatus); and claims 17-20 are directed to an article of
manufacture (i.e., a product); therefore, all pending claims are directed to one of the four
categories of invention.
Independent Claims
Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, independent claim 1 recites an abstract idea in the form of mental processes. A mental process is a process that “can be performed in the human mind, or by a human using a pen and paper” (MPEP§ 2106.04(a)(2)(III), paragraph 1). Examples of mental processes include “observations, evaluations, judgments, and opinions” (MPEP § 2106.04(a)(2)(III), paragraph 2).
The following limitations of claim 1 are mental processes:
choosing a state by selecting a source item and a target item; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for selecting an item is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
choosing an action based on the state using a policy, the policy depending on model parameters used by a plural-objective model to encode at least the source item, the plural-objective model being a model that is trained to promote plural objectives, [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for choosing an action is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
the action specifying whether the target item is selected because the target item matches the source item; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for the action specification is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
generating a reward based on the state and the action, the reward being based on, at least in part, whether the action is confirmed by at least one reference model; and [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for generating a reward is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
updating the model parameters used by the plural-objective model based on the reward. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for updating parameters is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
Therefore, the independent claims recite a judicial exception.
Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. The judicial exception recited in the above discussed claims is not integrated into a
practical application.
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. The claims do not include additional elements that are sufficient for the claims to
amount to significantly more than the judicial exception.
Claim 13
Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, independent claim 13 recites an abstract idea in the form of mental processes. A mental process is a process that “can be performed in the human mind, or by a human using a pen and paper” (MPEP§ 2106.04(a)(2)(III), paragraph 1). Examples of mental processes include “observations, evaluations, judgments, and opinions” (MPEP § 2106.04(a)(2)(III), paragraph 2).
The following limitations of claim 13 are mental processes:
generating output information based on at least one target item drawn from the set of one or more target items, [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for generating information is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
Therefore, the independent claims recite a judicial exception.
Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. The judicial exception recited in the above discussed claims is not integrated into a
practical application.
an instruction data store for storing computer-readable instructions; and [An instruction data store are components recited at a high level are construed as generic computer components used to implement the abstract idea. See MPEP 2106.05(f)(2). As such, the limitations do not integrate the abstract idea into a practical application. Nor to do they amount to significantly more.]
a processing system for executing the computer-readable instructions in the data store, to perform operations including: [A processing system are components recited at a high level are construed as generic computer components used to implement the abstract idea. See MPEP 2106.05(f)(2). As such, the limitations do not integrate the abstract idea into a practical application. Nor to do they amount to significantly more.]
receiving the input query; [receiving a query is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).]
using the plural-objective model to identify a set of one or more target items in response to the query, the plural-objective model being a model that is trained to promote plural objectives; and [Identify a target item are mere instructions to apply the abstract idea. Mere recitation that a judicial exception is to be performed using generic class of computer algorithms in their ordinary capacity, cannot meaningfully integrate the judicial exception into a practical application. See MPEP 2106.05(f). Additionally, this is a description of how the abstract idea is performed, using a plural-objective model. As such, this merely describes a technological environment. See MPEP 2106.05(h).]
the plural-objective model having model parameters that have been trained by reinforcement learning to identify target items based on queries that are relevant to the respective source items and which differ, at least in part, from other target items produced by a novelty-reference model that is different than the plural-objective model, the novelty-reference model being a model that serves as a reference for assessing novelty. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the plural-objective model have trained parameters.].
Therefore, under MPEP 2106.04(d), the additional elements of the claims do not integrate
the judicial exception into a practical application.
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. The claims do not include additional elements that are sufficient for the claims to
amount to significantly more than the judicial exception.
Additional elements that are mere instructions to apply an exception or merely generally
linking or generally linking the use of a judicial exception to a particular technological
environment or field of use do not constitute significantly more than a judicial exception under
MPEP§2106.05(I)(A). Since the additional elements in the independent claims are all are mere
instructions to apply an exception or are merely generally linking or generally linking the use of
a judicial exception to a particular technological environment or field of use, they do not
constitute significantly more than a judicial exception.
Therefore, the additional elements identified in the Step 2A Prong Two analysis do not
constitute significantly more than a judicial exception.
Therefore, the independent claims are not patent eligible.
Claim 17
Step 2A Prong One: Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, independent claim 17 recites an abstract idea in the form of mental processes. A mental process is a process that “can be performed in the human mind, or by a human using a pen and paper” (MPEP§ 2106.04(a)(2)(III), paragraph 1). Examples of mental processes include “observations, evaluations, judgments, and opinions” (MPEP § 2106.04(a)(2)(III), paragraph 2).
The following limitations of claim 17 are mental processes:
choosing a state by selecting a source item and a target item; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for selecting an item is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
choosing an action based on the state using a policy, the policy depending on model parameters used by a plural-objective model to encode at least the source item, the plural-objective model being a model that is trained to promote plural objectives, [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for choosing an action is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
the action specifying whether the target item is selected because the target item matches the source item; [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for the action specification is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
generating a reward based on, at least in part, the set the candidate target items and the relevance result; and [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for generating a reward is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
updating the model parameters used by the plural-objective model based on the reward. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for updating parameters is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
Therefore, the independent claims recite a judicial exception.
Step 2A Prong Two: Does the claim recite additional elements that integrate the judicial exception into a practical application?
No. The judicial exception recited in the above discussed claims is not integrated into a
practical application.
A computer-readable storage medium for storing computer-readable instructions, a processing system executing the computer-readable instructions to perform operations, the operations comprising each of: [A computer-readable storage medium are components recited at a high level are construed as generic computer components used to implement the abstract idea. See MPEP 2106.05(f)(2). As such, the limitations do not integrate the abstract idea into a practical application. Nor to do they amount to significantly more.]
receiving a set of candidate target items that a novelty-reference model generates based on the source item, the novelty-reference model being different than the plural-objective model, the novelty-reference model being a model that serves as a reference for assessing novelty; [receiving items is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).]
receiving a relevance result that a relevance-reference model generates based on a prompt, the relevance result specifying whether the source item is relevant to the target item, the relevance-reference model being a model that serves as a reference for assessing relevance, the relevance-reference model being different than the plural-objective model; [receiving a result is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).]
Therefore, under MPEP 2106.04(d), the additional elements of the claims do not integrate
the judicial exception into a practical application.
Step 2B: Does the claim recite additional elements that amount to significantly more than the judicial exception?
No. The claims do not include additional elements that are sufficient for the claims to
amount to significantly more than the judicial exception.
Additional elements that are mere instructions to apply an exception or merely generally
linking or generally linking the use of a judicial exception to a particular technological
environment or field of use do not constitute significantly more than a judicial exception under
MPEP§2106.05(I)(A). Since the additional elements in the independent claims are all are mere
instructions to apply an exception or are merely generally linking or generally linking the use of
a judicial exception to a particular technological environment or field of use, they do not
constitute significantly more than a judicial exception.
Therefore, the additional elements identified in the Step 2A Prong Two analysis do not
constitute significantly more than a judicial exception.
Therefore, the independent claims are not patent eligible.
Dependent Claims
The remaining dependent claims being rejected do not recite additional elements, whether considered individually or in combination, that are sufficient to integrate the judicial exception into a practical application or amount to significantly more than the judicial exception.
Claim 2
wherein the plural-objective model, at a start of the training, includes pre-trained parameters produced based on supervised training. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the plural-objective model includes pre-trained parameters.].
Claim 3
wherein the selecting of the source item includes sampling the source item from a data store of source items. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the selecting of the source item includes sampling the source item.].
Claim 4
wherein the selecting of the target item includes sampling the target item based on probability information produced by the plural-objective model based on the source item, the probability information describing likelihoods of different candidate items matching the source item. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions]
Claim 5
wherein the selecting of the target item includes sampling the target item from plural subsets of candidate target items produced by different item-selecting techniques, one of the techniques using the plural-objective model. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the selecting of the target item includes sampling the target item.].
Claim 6
wherein the generating of the reward includes receiving a set of candidate target items that a novelty-reference model generates based on the source item, and determining whether the target item is among the set of candidate target items, the novelty-reference model being different than the plural-objective model, the novelty-reference model being a model that serves as a reference for assessing novelty. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for generating a reward is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
Claim 7
wherein the novelty-reference model has been trained using supervised training based on a training set that specifies pairs of items that are considered associated and pairs of items that are considered non-associated, based on a specified standard of association. [Training a model are mere instructions to apply the abstract idea. Mere recitation that a judicial exception is to be performed using generic class of computer algorithms in their ordinary capacity, cannot meaningfully integrate the judicial exception into a practical application. See MPEP 2106.05(f)]
Claim 8
wherein the generating of the reward includes receiving a relevance result that a relevance-reference model generates based on a prompt, the relevance-reference model being different than the plural-objective model, the relevance-reference model being a model that serves as a reference for assessing relevance, [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the generating of the reward includes receiving a relevance result].
the prompt including a description of the source item and the target item and instructions as to a task that the relevance-reference model is being asked to perform, and [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the prompt includes descriptions and instructions.].
the relevance result indicating whether the target item is relevant to the source item. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions.]
Claim 9
wherein the relevance-reference model is a language model that autoregressively generates the relevance result. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the relevance-reference model is a language model.].
Claim 10
receiving a set of candidate target items that a novelty-reference model generates based on the source item, the novelty-reference model being a model that serves as a reference for assessing novelty; [receiving a set of items is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).]
receiving a relevance result that a relevance-reference model generates based on a prompt, the relevance result specifying whether the source item is relevant to the target item, the relevance-reference model being a model that serves as a basis for assessing relevance; and [receiving a result is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).]
generating the reward based on, at least in part, the set of candidate target items and the relevance result, [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for generating a reward is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
the novelty-reference model and the relevance-reference model being models that are different than the plural-objective model. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the novelty-reference model and the relevance-reference model are different than the plural-objective model].
Claim 11
wherein the plural-objective model includes a first encoder for mapping the source item into first encoder output information, and a second encoder for mapping the target item into second encoder output information, and logic for generating a score that expresses an extent to which the second encoder output information matches the first encoder output information, and [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the plural-objective model includes a first and second encoder, and logic].
wherein the updating of the model parameters includes updating the model parameters used by the first encoder and the second encoder. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely updating of the model parameters includes updating the model parameters used by the first and second encoder].
Claim 12
wherein the plural-objective model includes a first encoder for mapping the source item into first encoder output information, [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the plural-objective model includes a first encoder.].
wherein pre-generated second encoder output information associated with the target item is retrieved from a data store, [retrieving information is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).]
wherein the plural-objective model further includes logic for generating a score that expresses an extent to which the second encoder output information matches the first encoder output information, and [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the plural-objective model further includes logic for generating a score.].
wherein the updating of the model parameters includes updating the model parameters used by the first encoder, encoder output information pertaining to candidate target items remaining fixed. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the updating of the model parameters includes updating the model parameters used by the first encoder.].
Claim 14
wherein the using the plural-objective model comprises using the plural-objective model to generate first encoder output information based on the query, and comparing the first encoder output information with each of plural instances of second encoder output information associated with different respective target items. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the using the plural-objective model comprises generating and comparing first encoder output information].
Claim 15
wherein the reinforcement learning represents each state as a particular query and a particular target item, wherein an action associated with the state is an indication of whether the particular target item is selected because the particular target item matches the query. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the reinforcement learning represents each state as a query and target item.].
Claim 16
receiving a set of candidates target items that the novelty-reference model generates based on the particular query; [receiving a set is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).]
receiving a relevance result that a relevance-reference model generates based on a prompt, the relevance result specifying whether the particular query is relevant to the particular target item, the relevance-reference model being a model that serves as a reference for assessing relevance; and [receiving a result is sending data, which is insignificant, extra-solution activity. See MPEP 2106.05(g). Transmitting data is well-understood, routine, and conventional. See MPEP 2106.05(d)(II)(i).]
generating the reward based on, at least in part, the set of candidate target items and the relevance result. [This is a mental process that can be performed by observations, evaluations, judgments, and opinions. No specific methodology for generating the reward is recited in the claim; therefore, it broadly encompasses processing that can be performed as a mental process.]
Claim 18
wherein the plural-objective model, at start of training, includes pre-trained parameters produced based on supervised training. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the plural-objective model includes pre-trained parameters.].
Claim 19
wherein the novelty-reference model has been trained using supervised training based on a training set that specifies pairs of items that are considered associated and pairs of items that are considered non-associated, based on a specified standard of association. [Training a model are mere instructions to apply the abstract idea. Mere recitation that a judicial exception is to be performed using generic class of computer algorithms in their ordinary capacity, cannot meaningfully integrate the judicial exception into a practical application. See MPEP 2106.05(f)]
Claim 20
wherein the selecting of the source item includes sampling the source item from a data store of source items, and [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the selecting of the source item includes sampling the source item.].
wherein the selecting of the target item includes sampling the target item based on probability information produced by the plural-objective model based on the source item, the probability information describing likelihoods of different candidate items matching the source item. [This additional element does no more than generally link the use of a judicial exception to a particular technological environment or field of use (MPEP § 2106.05(h)). This element merely indicates a field of use or technological environment in which a judicial exception is applied, namely the selecting of the target item includes sampling the target item.].
The prior art used for rejections are provided below:
Choosing the Best of Both Worlds: Diverse and Novel Recommendations through Multi-Objective Reinforcement Learning (October 28, 2021) to Stamenković et al. (hereinafter Stamenković)
Enhancing Generative Retrieval with Reinforcement Learning from Relevance Feedback (December 6, 2023) to Zhou et al. (hereinafter Zhou)
Dense Passage Retrieval for Open-Domain Question Answering (November 16, 2020) to Karpukhin et al. (hereinafter Karpukhin)
Selective Weak Supervision for Neural Information Retrieval (January 28, 2020) to Zhang et al. (hereinafter Zhang)
Perspectives on Large Language Models for Relevance Judgment (November 18, 2023) to Faggioli et al. (hereinafter Faggioli)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1-6, 11, 13, 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Stamenković
Per claim 1, Zhang discloses: A method for training a machine-trained model, comprising: [Zhang, pg. 1, 1 Introduction "This data selector is connected to the target neural ranker using policy gradients... The learning of the two parties is conducted iteratively in ReInfoSelect's stochastic process." (note: Zhang trains a machine-learned data selector (its state and action networks) iteratively using policy gradients)]
choosing a state by selecting a source item and a target item; [Zhang, pg. 3, 3.2 Reinforcement Data Selection "for the i-th weak supervision pair bi = (ai,di)... The state si for the i-th pair include three continuous vectors: anchor state
s
i
a
, document state
s
i
d
, and anchor-document interaction state
s
i
a
d
." (note: the state si is created directly and non-circularly from the pair (ai, di), the anchor (source item) and the document (target item), such that the state representation itself is created from selecting both items together)]
the action specifying whether the target item is selected because the target item matches the source item; [Zhang, pg. 4, 3.2 Reinforcement Data Selection "The action decides whether to use the anchor-document pair (1) or not (0) as a weak supervision signal. The action on the i-th a-d pair is calculated as Actioni = argmax0,1 π(si)."; pg. 2, 3.1 Preliminary “matches the query and document in the n-gram Conv-KNRM embedding space using matching kernels [45].” (note: the binary action operates on the same pair (ai, di) already designated as the state in the previous limitation, a determination of whether that already-paired document is confirmed as a match for that specific anchor. Conv-KNRM explicitly stated to match the query and document via kernels.)]
generating a reward based on the state and the action, the reward being based on, at least in part, whether the action is confirmed by at least one reference model;
[Zhang, pg. 4, 3.2 Reinforcement Data Selection "
PNG
media_image1.png
63
609
media_image1.png
Greyscale
NDCG evaluates the neural ranker's accuracy on the validation part of the target ranking benchmarks." (note: the reward is computed from a separate model (the neural ranker f) evaluated against validation relevance labels, functioning as the reference model whose confirmation (improved NDCG) determines the reward.)]
updating the model parameters used by the plural-objective model based on the reward. [Zhang, pg. 4, 3.2 Reinforcement Data Selection "
PNG
media_image2.png
82
672
media_image2.png
Greyscale
" (note: this is a policy-gradient update of the state/action network's parameters using the reward computed")]
Zhang does not fully disclose, but Zhang with Stamenković does teach:
choosing an action based on the state using a policy, the policy depending on model parameters used by a plural-objective model to encode at least the source item, the plural-objective model being a model that is trained to promote plural objectives; [Stamenković, pg. 1, Abstract "The proposed SMORL agent augments standard recommendation models with additional RL layers that enforce it to simultaneously satisfy three principal objectives: accuracy, diversity, and novelty of recommendations." (note: Zhang's state/action network is dependent on model parameters that encode the source item, and Zhang optimizes a single objective (NDCG) s. Stamenković states a model trained via RL to promote three objectives (accuracy, diversity, novelty). Combining Stamenković's plural-objective training with Zhang's parameter-dependent, source item encoding policy structure is this limitation.)]
Zhang and Stamenković are analogous art because they are from the same field of endeavor of reinforcement-learning-based training of models that identify, select, or confirm target items relative to source items using retrieval- or ranking-quality metrics as the reward signal. They are further reasonably pertinent to the same problem of training such models using reinforcement learning because the target task cannot be trained with ordinary differentiable methods.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to substitute Stamenković’s multi-objective reward structure for Zhang’s single-objective NDCG reward within the existing state/action architecture.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”].
Per claim 2, Zhang-Stamenković discloses claim 1.
Zhang further teaches the plural-objective model, at a start of the training, includes pre-trained parameters produced based on supervised training. [Zhang, pg. 5, 4 Experimental Methodology "There are three steps for training with the ReInfoSelect: warm up, reinforce training with anchor data, and adapting to the ranking benchmark... The warm up stage first trains the state and action network using the discriminator setup [26]. Then in the reinforce stage, ReInfoSelect's networks are initialized (warmed up) by the learned discriminator weights." (note: Zhang discloses a two-phase training pipeline for the state/action network, a warm up phase precedes the reinforcement learning phase, with the reinforcement learning phase explicitly initialized from the pre-trained supervised weights)]
Per claim 3, Zhang-Stamenković discloses claim 1.
Zhang further teaches the selecting of the source item includes sampling the source item from a data store of source items. [Zhang, pg. 4, 4 Experimental Methodology "About 100K anchors (from total 6 million collected) and their linked documents are randomly sampled as the weak supervision dataset." (note: Zhang discloses a data store (6 million collected anchors) from which source items (anchors) are randomly sampled)]
Per claim 4, Zhang-Stamenković discloses claim 1.
Zhang does not fully disclose, but with Stamenković does teach: the selecting of the target item includes sampling the target item based on probability information produced by the plural-objective model based on the source item, the probability information describing likelihoods of different candidate items matching the source item. [Stamenković, pg. 3, 4 Model And Training "we cast the task of next item recommendation as a (self-supervised) multi-class classification problem and build a sequential model that receives user-item interaction sequence x1:t... and generates n classification logits yt+1 = [y1, y2, ..., yn] ∈ Rn"; pg. 4, 4 MODEL AND TRAINING “we train our recommendation model by optimizing the cross-entropy loss Ls based on the logits yt+1” (note: the quantity that is interpreted as this limitation is softmax(yt+1), the softmax normalized version of Stamenković’s classification logits. Stamenković states that these logits are used for a cross-entropy loss. The softmax step is a necessary, standard part of cross-entropy loss over multi-class logits.)]
Zhang and Stamenković are analogous art because they are from the same field of endeavor of reinforcement-learning-based training of models that identify, select, or confirm target items relative to source items using retrieval- or ranking-quality metrics as the reward signal. They are further reasonably pertinent to the same problem of training such models using reinforcement learning because the target task cannot be trained with ordinary differentiable methods.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to substitute Stamenković’s multi-objective reward structure for Zhang’s single-objective NDCG reward within the existing state/action architecture.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”].
Per claim 5, Zhang-Stamenković discloses claim 1.
Zhang does not fully disclose, but with Stamenković does teach: the selecting of the target item includes sampling the target item from plural subsets of candidate target items produced by different item-selecting techniques, one of the techniques using the plural-objective model. [Stamenković, pg. 3, 4 MODEL AND TRAINING “When generating recommendations, we still return the top-k items from the supervised head”; pg. 5, 4.3 Reinforcing Novelty “
PNG
media_image3.png
89
704
media_image3.png
Greyscale
”. (note: there are two named subsets, the first is the top-k items returned by the plural-objective model’s supervised head, the second is the top-x% most popular items, which is produced by a separate popularity-based mechanism)]
Zhang and Stamenković are analogous art because they are from the same field of endeavor of reinforcement-learning-based training of models that identify, select, or confirm target items relative to source items using retrieval- or ranking-quality metrics as the reward signal. They are further reasonably pertinent to the same problem of training such models using reinforcement learning because the target task cannot be trained with ordinary differentiable methods.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to substitute Stamenković’s multi-objective reward structure for Zhang’s single-objective NDCG reward within the existing state/action architecture.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”].
Per claim 6, Zhang-Stamenković discloses claim 1.
Zhang does not fully disclose, but with Stamenković does teach: the generating of the reward includes receiving a set of candidate target items that a novelty-reference model generates based on the source item, and determining whether the target item is among the set of candidate target items, the novelty-reference model being different than the plural-objective model, the novelty-reference model being a model that serves as a reference for assessing novelty [Stamenković, pg. 5, 4.3 Reinforcing Novelty “
PNG
media_image3.png
89
704
media_image3.png
Greyscale
”. (note: this shows the novelty reward is computed by checking membership of the selected item pt in a reference set (the top x% most popular items, a set generated with respect to the given interaction context (source item)), and assigning a binary reward based on that membership. The popularity ranking mechanism is the novelty-reference model and the reward computation is the membership-based reward logic.)].
Zhang and Stamenković are analogous art because they are from the same field of endeavor of reinforcement-learning-based training of models that identify, select, or confirm target items relative to source items using retrieval- or ranking-quality metrics as the reward signal. They are further reasonably pertinent to the same problem of training such models using reinforcement learning because the target task cannot be trained with ordinary differentiable methods.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to substitute Stamenković’s multi-objective reward structure for Zhang’s single-objective NDCG reward within the existing state/action architecture.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”].
Per claim 11, Zhang-Stamenković discloses claim 1.
Zhang further teaches the plural-objective model includes a first encoder for mapping the source item into first encoder output information, and a second encoder for mapping the target item into second encoder output information, [Zhang, pg. 4, 3.2 Reinforcement Data Selection "The anchor state and document state representations use standard convolutional neural networks on their word embeddings:
s
i
a
= CNNa(ai);
s
i
d
= CNNd(di)." (note: Zhang teaches two separate encoders, CNNa for the source item (anchor) and CNNd for the target item (document))] and logic for generating a score that expresses an extent to which the second encoder output information matches the first encoder output information [Zhang, pg. 4, 3.2 Reinforcement Data Selection "The anchor-document state representation uses the ranking feature Φ(Mai,di) from Conv-KNRM (Eqn. 4):
s
i
a
d
= Φ(Mai,di)." (note: the anchor-document state (score) is computed via a ranking feature built from a similarity translation matrix between the anchor's and document's embeddings)], and
wherein the updating of the model parameters includes updating the model parameters used by the first encoder and the second encoder. [Zhang, pg. 4, 3.2 Reinforcement Data Selection "Note that the parameters are not shared with the neural ranker." (note: CNNa, CNNd, and the interaction feature all belong to the state network being trained, not the separate downstream ranker, and the policy-gradient update applies to the entire state network's parameter set, therefore both encoders are updated together by the same gradient step.)]
Per claim 13, Zhang discloses: A computing system for processing an input query using a machine-trained model, comprising [Zhang, pg. 5, 4 Experimental Methodology “All our neural models are implemented with PyTorch. All models are trained with a single GeForce GTX TITAN GPU”. (note: Zhang’s implementation includes a computing system executing a trained model to process queries)]:
an instruction data store for storing computer-readable instructions; and [Zhang, pg. 1, footnote "All our codes, data, and results are available at https://github.com/thunlp/ReInfoSelect." (note: Zhang's method exists as stored computer-readable instructions)]
a processing system for executing the computer-readable instructions in the data store, to perform operations including: [Zhang, pg. 4, 4 Experimental Methodology ("Implementation Details")]: "The neural ranking models are updated with one gradient step per batch, while the data selector is updated once every 4 batches (T=4)." (note: this shows execution parameters (gradient steps per batch, update cadence), which confirm actual execution on computing hardware)]
receiving the input query; [Zhang, pg. 3, 3.2 Reinforcement Data Selection “for the i-th weak supervision pair, bi = (ai, di)”. (note: the anchor ai (source item/query) is received as input to the state pipeline)]
generating output information based on at least one target item drawn from the set of one or more target items, [Zhang, pg. 4, 3.3 Neural Ranker Training with ReInfoSelect "ReInfoSelect first selects the anchor-document pairs using its state and action networks, and then stochastically trains the neural ranker using the selected pairs." (note: output information is generated from the target items.)]
Zhang does not fully disclose, but Zhang with Stamenković does teach: using the plural-objective model to identify a set of one or more target items in response to the query, the plural-objective model being a model that is trained to promote plural objectives [Stamenković, pg. 1, Abstract "The proposed SMORL agent augments standard recommendation models with additional RL layers that enforce it to simultaneously satisfy three principal objectives: accuracy, diversity, and novelty of recommendations." (note: Stamenković discloses the plural-objectives training characterization)]; and
the plural-objective model having model parameters that have been trained by reinforcement learning to identify target items based on queries that are relevant to the respective source items and which differ, at least in part, from other target items produced by a novelty-reference model that is different than the plural-objective model, the novelty-reference model being a model that serves as a reference for assessing novelty. [Stamenković, pg. 5, 4.3 Reinforcing Novelty “
PNG
media_image3.png
89
704
media_image3.png
Greyscale
”; pg. 3, 2 RELATED WORK “SMORL significantly increases diversity and slightly improves the accuracy”; pg. 6, 5.2 Performance Comparison (RQ1) “we consistently outperform the corresponding baselines across all metrics” (note: rnov equation is the training mechanism showing model parameters are updated using a reward that is zero when the selected item is within the novelty-reference model’s output (top x% of most popular items) and positive otherwise, therefore directly training the model to identify items that differ from what the novelty-reference model would produce).)
Zhang and Stamenković are analogous art because they are from the same field of endeavor of reinforcement-learning-based training of models that identify, select, or confirm target items relative to source items using retrieval- or ranking-quality metrics as the reward signal. They are further reasonably pertinent to the same problem of training such models using reinforcement learning because the target task cannot be trained with ordinary differentiable methods.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to substitute Stamenković’s multi-objective reward structure for Zhang’s single-objective NDCG reward within the existing state/action architecture.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”].
Per claim 15, Zhang-Stamenković discloses claim 13.
Zhang further teaches the reinforcement learning represents each state as a particular query and a particular target item, wherein an action associated with the state is an indication of whether the particular target item is selected because the particular target item matches the query. [Zhang, pg. 3, 3.2 Reinforcement Data Selection "for the i-th weak supervision pair bi = (ai,di)... The state si for the i-th pair include three continuous vectors: anchor state
s
i
a
, document state
s
i
d
, and anchor-document interaction state
s
i
a
d
." ; pg. 4, 3.2 Reinforcement Data Selection "The action decides whether to use the anchor-document pair (1) or not (0) as a weak supervision signal. The action on the i-th a-d pair is calculated as Actioni = argmax0,1 π(si)." (note: the state si is created directly from the pair (ai, di), the anchor (query/source item) and the document (target item), and the binary action is the matching indication for that same pair)]
Claim(s) 7, 12, and 14 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Stamenković, in further view of Karpukhin.
Per claim 7, Zhang-Stamenković discloses claim 6.
Zhang does not fully disclose, but with Stamenković-Karpukhin does teach: the novelty-reference model has been trained using supervised training based on a training set [Karpukhin, pg. 6771, 3.2 Training “
PNG
media_image4.png
68
627
media_image4.png
Greyscale
”. (note: Karpukhin defines its training data D as a collection of instances, each comprising a source item and its associated and non-associated target items)] that specifies pairs of items that are considered associated [Karpukhin, pg. 6771, 3.2 Training “Each instance contains one question qi and one relevant (positive) passage
p
i
+
”. (note: the positive passage
p
i
+
is designated as relevant to the question qi within each training instance, which is pairs of items that are considered associated)] and pairs of items that are considered non-associated [Karpukhin, pg. 6771, 3.2 Training “…along with n irrelevant (negative) passages pi,j”. (note: the negative passages pi,j are designated irrelevant within each training instance, which is pairs of items that are considered non-associated)], based on a specified standard of association [Karpukhin, pg. 6771, 3.2 Training “passages relevant to a question may be given in a QA dataset, or can be found using the answer. All other passages in the collection, while not specified explicitly, can be viewed as irrelevant by default.”. (note: this shows the standard used for labeling passages positive versus negative by default, answer containment or relevance to a question. This is a specified standard of association.)].
Zhang, Stamenković, and Karpukhin are analogous art because they are from the same field of endeavor of models that generate and compare source item and target item representations for ranking purposes.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to substitute Stamenković’s multi-objective reward structure for Zhang’s single-objective NDCG reward within the existing state/action architecture, and the explicit passage index architecture as taught by Karpukhin to implement Stamenković's freezing technique.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”]. Karpukhin’s indexed dual-encoder approach improves retrieval effectiveness. [Karpukhin, pg. 6770, 1 Introduction “we demonstrate that with the proper training setup, simply fine-tuning the question and passage encoders on existing question-passage pairs is sufficient to greatly outperform BM25.”.
Per claim 12, Zhang-Stamenković discloses claim 1.
Zhang does not fully disclose, but with Stamenković-Karpukhin does teach: the plural-objective model includes a first encoder for mapping the source item into first encoder output information, [Stamenković, pg. 3, 4 MODEL AND TRAINING "Typically one can use a generative sequence model G(·) to map the input sequence into a hidden state st = G(x1:t). This serves as a general encoder function." (note: G is stated as an encoder function mapping the source item to encoder output information st)]
wherein pre-generated second encoder output information associated with the target item is retrieved from a data store, [Karpukhin, pg. 6770, 3.1 Overview "Our dense passage retriever (DPR) uses a dense encoder EP (·) which maps any text passage to a d-dimensional real-valued vectors and builds an index for all the M passages that we will use for retrieval." (note: Karpukhin shows building an index (data store) of pre-generated encoder outputs for retrieval)]
wherein the plural-objective model further includes logic for generating a score that expresses an extent to which the second encoder output information matches the first encoder output information [Stamenković, pg. 4, 4.2 Reinforcing Diversity “
PNG
media_image5.png
77
621
media_image5.png
Greyscale
”. (note: this similarity computation generates a score expressing the match between the fixed reference embedding (second encoder output information) and the source-associated embedding (first encoder output information))], and
wherein the updating of the model parameters includes updating the model parameters used by the first encoder, encoder output information pertaining to candidate target items remaining fixed. [Stamenković, pg. 4, 4.2 Reinforcing Diversity "We then freeze the weights of Ediv to stop further updates of the parameters" ; pg. 5, Algorithm 1 “
PNG
media_image6.png
59
441
media_image6.png
Greyscale
” (note: Ediv (pertaining to candidate target items) is fixed, while gradient updates apply only to Θ, the parameters of the base model G (the first encoder)]
Zhang, Stamenković, and Karpukhin are analogous art because they are from the same field of endeavor of models that generate and compare source item and target item representations for ranking purposes.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to substitute Stamenković’s multi-objective reward structure for Zhang’s single-objective NDCG reward within the existing state/action architecture, and the explicit passage index architecture as taught by Karpukhin to implement Stamenković's freezing technique.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”]. Karpukhin’s indexed dual-encoder approach improves retrieval effectiveness. [Karpukhin, pg. 6770, 1 Introduction “we demonstrate that with the proper training setup, simply fine-tuning the question and passage encoders on existing question-passage pairs is sufficient to greatly outperform BM25.”.
Per claim 14, Zhang-Stamenković discloses claim 13.
Zhang does not fully disclose, but with Stamenković-Karpukhin does teach: the using the plural-objective model comprises using the plural-objective model to generate first encoder output information based on the query, and comparing the first encoder output information with each of plural instances of second encoder output information associated with different respective target items [Karpukhin, pg. 6771, 3.1 Overview “At run-time, DPR applies a different encoder EQ(⋅) that maps the input question to a d-dimensional vector, and retrieves k passages of which vectors are the closest to the question vector”. (note: this shows generating first encoder output (the question vector via EQ) based on the query, and then comparing that output against plural pre-stored instances of second encoder output (the pre-indexed passage vectors, generated offline by EP) to retrieve the closest matches.)].
Zhang, Stamenković, and Karpukhin are analogous art because they are from the same field of endeavor of models that generate and compare source item and target item representations for ranking purposes.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to substitute Stamenković’s multi-objective reward structure for Zhang’s single-objective NDCG reward within the existing state/action architecture, and the explicit passage index architecture as taught by Karpukhin to implement Stamenković's freezing technique.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”]. Karpukhin’s indexed dual-encoder approach improves retrieval effectiveness. [Karpukhin, pg. 6770, 1 Introduction “we demonstrate that with the proper training setup, simply fine-tuning the question and passage encoders on existing question-passage pairs is sufficient to greatly outperform BM25.”.
Claim(s) 8 and 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Stamenković, in further view of Zhou and Faggioli.
Per claim 8, Zhang-Stamenković discloses claim 1.
Zhang does not fully disclose, but with Stamenković-Zhou-Faggioli does teach: the generating of the reward includes receiving a relevance result that a relevance-reference model generates based on a prompt, the relevance-reference model being different than the plural-objective model, the relevance-reference model being a model that serves as a reference for assessing relevance [Zhou, pg. 12484, 4.2 Relevance Reward Model Training “For assessing contextual dependency, we employ the large language model (LLM), LLaMA-13b (Touvron et al., 2023), to compute the generation likelihood of the document given a specific query. We assume that the generation probability can serve as an indicative metric of how well the document fulfills the implicit demands enclosed in the query.”. (note: this shows a different model generating a relevance result from a prompt, which serves a relevance result); pg. 12486, 5.3 Implementation Details “Our SFT model utilizes the T5-base pretrained model (Raffel et al., 2020) as the backbone” (note: this shows the model being trained is T5-based, therefore distinct from LLaMA-13b (plural-objective model))],
the prompt including a description of the source item and the target item and instructions as to a task that the relevance-reference model is being asked to perform [Zhou, pg. 12484, 4.2 Relevance Reward Model Training “For assessing contextual dependency, we employ the large language model (LLM), LLaMA-13b (Touvron et al., 2023), to compute the generation likelihood of the document given a specific query.” (note: this shows the input includes both the query (source item) and document (target item);
Faggioli, pg. 18, 5.1 Methodology “Instruction: You are an expert assessor making TREC relevance judgments. You will be given a TREC topic and a portion of a document. If any part of the document is relevant to the topic, answer “Yes”. If not, answer “No”. Remember that the TREC relevance condition states that a document is relevant to a topic if it contains information that is helpful in satisfying the user’s information need described by the topic. A document is judged relevant if it contains information that is on-topic and of potential value to the user. Topic: {topic} Document: {document} Relevant?” (note: this is an explicit prompt template used to bring forth relevance judgement, and states instructions for a task))], and
the relevance result indicating whether the target item is relevant to the source item [Zhou, pg. 12484, 4.2 Relevance Reward Model Training “We assume that the generation probability can serve as an indicative metric of how well the document fulfills the implicit demands enclosed in the query.”. (note: this shows the model’s output indicates whether the target item (document) is relevant to the source item (query))].
Zhang, Stamenković, Zhou, and Faggili are analogous art because they are from the same field of endeavor of evaluating relevance determinations between a source item and a target item. They are further reasonably pertinent to the same problem of obtaining a reliable relevance signal without relying only on human labeled supervision
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to incorporate Stamenković’s multi-objective reward structure, Zhou’s LLM relevance scoring mechanism, and Faggioli’s instruction prompt design into Zhang’s existing state and action framework.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”]. Incorporating Zhou’s relevance signal leads to improved performance [Zhou, pg. 12487, 6.2 Ablation Study on Relevance Annotators “the inclusion of the LLM annotator yields performance enhancements, indicating that leveraging the zero-shot capability of large language models can further augment the model’s ability to capture contextual dependencies.”]. Including the prompt with explicit task instructions as taught by Faggioli constrains the model’s output to usable and parseable relevance information [Faggioli, pg. 18, 5.1 Methodology “the prompts and the setting Temperature = 0 were sufficient to constrain the model to emit only the relevance grades requested in the prompt”].
Per claim 9, Zhang-Stamenković-Zhou-Faggioli discloses claim 8.
Zhang does not fully disclose, but with Stamenković-Zhou does teach: the relevance-reference model is a language model that autoregressively generates the relevance result [Zhou, pg. 12484, 4.2 Relevance Reward Model Training “we employ the large language model (LLM), LLaMA-13b (Touvron et al., 2023), to compute the generation likelihood of the document given a specific query.”. (note: LLaMA-13b is an autoregressive transformer language model and is used to determine relevance)].
Zhang, Stamenković, Zhou, and Faggili are analogous art because they are from the same field of endeavor of evaluating relevance determinations between a source item and a target item. They are further reasonably pertinent to the same problem of obtaining a reliable relevance signal without relying only on human labeled supervision
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to incorporate Stamenković’s multi-objective reward structure, Zhou’s LLM relevance scoring mechanism, and Faggioli’s instruction prompt design into Zhang’s existing state and action framework.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”]. Incorporating Zhou’s relevance signal leads to improved performance [Zhou, pg. 12487, 6.2 Ablation Study on Relevance Annotators “the inclusion of the LLM annotator yields performance enhancements, indicating that leveraging the zero-shot capability of large language models can further augment the model’s ability to capture contextual dependencies.”]. Including the prompt with explicit task instructions as taught by Faggioli constrains the model’s output to usable and parseable relevance information [Faggioli, pg. 18, 5.1 Methodology “the prompts and the setting Temperature = 0 were sufficient to constrain the model to emit only the relevance grades requested in the prompt”].
Claim(s) 10, 16-18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Stamenković, in further view of Zhou.
Per claim 10, Zhang-Stamenković discloses claim 1.
Zhang does not fully disclose, but with Stamenković-Zhou does teach: the generating of the reward includes: receiving a set of candidate target items that a novelty-reference model generates based on the source item, the novelty-reference model being a model that serves as a reference for assessing novelty [Stamenković, pg. 5, 4.3 Reinforcing Novelty “
PNG
media_image3.png
89
704
media_image3.png
Greyscale
”. (note: this shows receiving a reference set (top x% popular items, computed from the training population) to determine the target item’s novelty status)];
receiving a relevance result that a relevance-reference model generates based on a prompt, the relevance result specifying whether the source item is relevant to the target item, the relevance-reference model being a model that serves as a basis for assessing relevance [Zhou, pg. 12484, 4.2 Relevance Reward Model Training “we employ the large language model (LLM), LLaMA-13b (Touvron et al., 2023), to compute the generation likelihood of the document given a specific query.”. (note: this shows generating a relevance result from an LLM reference model distinct from the retrieval model based on a query document input)]; and
generating the reward based on, at least in part, the set of candidate target items and the relevance result [Stamenković, pg. 5, Algorithm 1 “
PNG
media_image7.png
69
608
media_image7.png
Greyscale
”. (note: Stamenković generates a single reward value from multiple separately computed reward components. It explicitly stacks racc (an accuracy/relevance signal), rdiv, and rnov (novelty signal) into one vector rt. This states the technique of generating the reward based on, at least in part, the set of candidate target items (rnov, the novelty reference) and the relevance result (racc). This results in a single-generated reward incorporating both signal types. As per the previous limitation, Zhou teaches the relevance type component should be produced by an LLM configured to provide reference values. Therefore, by substituting Zhou’s LLM generated relevance result into Stamenković’s ‘stack(…)’ reward combination mechanism in place of or alongside racc arrives directly at this limitation, a single generated reward based on both the novelty reference candidate set and the LLM relevance result. This also matches Stamenković’s own disclosed technique of combining a novelty-based reward component with a relevance-based reward component into one value.)],
the novelty-reference model and the relevance-reference model being models that are different than the plural-objective model [Stamenković, pg. 4, 4.2 Reinforcing Diversity “we first train a GRU4Rec model [14], and save the embedding layer Ediv. We then freeze the weights of Ediv to stop further updates of the parameters… We do not use the embedding of a model that is currently trained for calculation 𝑟div.”. (note: this shows that the reference embedding model (Ediv, used for novelty reward), is a separately trained frozen model, not the model currently being trained (the plural-objective model/base model G);
Zhou, pg. 12484, 4.2 Relevance Reward Model Training “we employ the large language model (LLM), LLaMA-13b (Touvron et al., 2023)”; pg. 12486, 5.3 Implementation Details “Our SFT model utilizes the T5-base pre trained model (Raffel et al., 2020) as the backbone”. (note: Zhou’s GenRRL’s own retrieval model being trained is a T5-base backbone, while a separate, distinct LLM (LLaMa-13b) is used for relevance. This shows the relevance-reference model is different from the plural-objective model.)].
Zhang, Stamenković, and Zhou are analogous art because they are from the same field of endeavor of training models using reinforcement learning to select or evaluate items relative to a source item, using reward signals obtained from separate reference models. They are further reasonably pertinent to the same problem of generating a reliable training signal when direct labeled supervision for the objective being optimized is limited or unavailable.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to include Stamenković’s multi-objective reward structure and Zhou’s LLM relevance signal into Zhang’s existing state and action framework.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”]. Incorporating Zhou’s relevance signal leads to improved performance [Zhou, pg. 12487, 6.2 Ablation Study on Relevance Annotators “the inclusion of the LLM annotator yields performance enhancements, indicating that leveraging the zero-shot capability of large language models can further augment the model’s ability to capture contextual dependencies.”].
Per claim 16, Zhang-Stamenković-Zhou discloses claim 15.
Zhang does not fully disclose, but with Stamenković-Zhou does teach: the reinforcement learning produces the model parameters based on a reward that is generated by: receiving a set of candidates target items that the novelty-reference model generates based on the particular query [Stamenković, pg. 5, 4.3 Reinforcing Novelty “
PNG
media_image3.png
89
704
media_image3.png
Greyscale
”. (note: the reference set (top x% of most popular items, a distribution computed from the training population and functioning as the novelty-reference model) is used to determine whether the top predicted item pt (the candidate produced in response to the given user state (query)) falls within that reference set. This is the receiving and use of a set of candidate target items that the novelty-reference model generates based on the particular query. The state st (and pt, the prediction generated from it) is derived from the particular query under evaluation, therefore the reference set is generated with respect to that specific query.)];
receiving a relevance result that a relevance-reference model generates based on a prompt, the relevance result specifying whether the particular query is relevant to the particular target item, the relevance-reference model being a model that serves as a reference for assessing relevance [Zhou, pg. 12484, 4.2 Relevance Reward Model Training “we employ the large language model (LLM), LLaMA-13b (Touvron et al., 2023), to compute the generation likelihood of the document given a specific query. We assume that the generation probability can serve as an indicative metric of how well the document fulfills the implicit demands enclosed in the query.” (note: GenRRL uses LLaMa-13b, an LLM distinct from the retrieval model being trained, as a relevance indicator. It receives a prompt comprising the query (particular query) and candidate document (particular target item) and returns a generation likelihood value which is a relevance result specifying whether the query is relevant to the target item. LLaMA-13b’s function in this pipeline is serving as a reference for assessing relevance.)]; and
generating the reward based on, at least in part, the set of candidate target items and the relevance result [Stamenković, pg. 5, Algorithm 1 “
PNG
media_image7.png
69
608
media_image7.png
Greyscale
”. (note: this shows how Stamenković generates a single reward by stacking multiple separately computed reward components (including rnov which is the candidate set/novelty signal) into one vector, then combining that vector by scalarization weight w into a single scalar used for training. Combined with Zhou’s teaching that the relevance type component of such a combined reward should be produced by an LLM oracle, a POSITA substituting Zhou’s LLM generated relevance results into Stamenković stack(…) combination mechanism arrives at this limitation, a single generated reward based on both the novelty-reference candidate set and the LLM relevance result.)].
Zhang, Stamenković, and Zhou are analogous art because they are from the same field of endeavor of training models using reinforcement learning to select or evaluate items relative to a source item, using reward signals obtained from separate reference models. They are further reasonably pertinent to the same problem of generating a reliable training signal when direct labeled supervision for the objective being optimized is limited or unavailable.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to include Stamenković’s multi-objective reward structure and Zhou’s LLM relevance signal into Zhang’s existing state and action framework.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”]. Incorporating Zhou’s relevance signal leads to improved performance [Zhou, pg. 12487, 6.2 Ablation Study on Relevance Annotators “the inclusion of the LLM annotator yields performance enhancements, indicating that leveraging the zero-shot capability of large language models can further augment the model’s ability to capture contextual dependencies.”].
Per claim 17, Zhang discloses: A computer-readable storage medium for storing computer-readable instructions, a processing system executing the computer-readable instructions to perform operations, the operations comprising each of [Zhang, pg. 1, footnote "All our codes, data, and results are available at https://github.com/thunlp/ReInfoSelect." (note: Zhang's method exists as stored computer-readable instructions)]
choosing a state by selecting a source item and a target item; [Zhang, pg. 3, 3.2 Reinforcement Data Selection "for the i-th weak supervision pair bi = (ai,di)... The state si for the i-th pair include three continuous vectors: anchor state
s
i
a
, document state
s
i
d
, and anchor-document interaction state
s
i
a
d
." (note: the state si is created directly from the pair (ai, di), the anchor (source item) and the document (target item), such that the state representation itself is created from selecting both items together)]
the action specifying whether the target item is selected because the target item matches the source item; [Zhang, pg. 4, 3.2 Reinforcement Data Selection "The action decides whether to use the anchor-document pair (1) or not (0) as a weak supervision signal. The action on the i-th a-d pair is calculated as Actioni = argmax0,1 π(si)."; pg. 2, 3.1 Preliminary “matches the query and document in the n-gram Conv-KNRM embedding space using matching kernels [45].” (note: the binary action operates on the same pair (ai, di) already designated as the state in the previous limitation, a determination of whether that already-paired document is confirmed as a match for that specific anchor. Conv-KNRM explicitly stated to match the query and document via kernels.)]
updating the model parameters used by the plural-objective model based on the reward. [Zhang, pg. 4, 3.2 Reinforcement Data Selection "
PNG
media_image2.png
82
672
media_image2.png
Greyscale
" (note: this is a policy-gradient update of the state/action network's parameters using the reward computed")]
Zhang does not fully disclose, but Zhang with Stamenković does teach:
choosing an action based on the state using a policy, the policy depending on model parameters used by a plural-objective model to encode at least the source item, the plural-objective model being a model that is trained to promote plural objectives; [Stamenković, pg. 1, Abstract "The proposed SMORL agent augments standard recommendation models with additional RL layers that enforce it to simultaneously satisfy three principal objectives: accuracy, diversity, and novelty of recommendations." (note: Zhang's state/action network is dependent on model parameters that encode the source item, and Zhang optimizes a single objective (NDCG) s. Stamenković states a model trained via RL to promote three objectives (accuracy, diversity, novelty). Combining Stamenković's plural-objective training with Zhang's parameter-dependent, source item encoding policy structure is this limitation.)]
receiving a set of candidate target items that a novelty-reference model generates based on the source item, the novelty-reference model being different than the plural-objective model, the novelty-reference model being a model that serves as a reference for assessing novelty [Stamenković, pg. 5, 4.3 Reinforcing Novelty “
PNG
media_image3.png
89
704
media_image3.png
Greyscale
”. (note: the popularity ranked reference set functions as the received candidate set from a novelty-reference model)];
generating a reward based on, at least in part, the set the candidate target items and the relevance result [Stamenković, pg. 5, Algorithm 1 “
PNG
media_image7.png
69
608
media_image7.png
Greyscale
”. (note: Stamenković generates a single reward value from multiple separately computed reward components. It explicitly stacks racc (an accuracy/relevance signal), rdiv, and rnov (novelty signal) into one vector rt. This states the technique of generating the reward based on, at least in part, the set of candidate target items (rnov, the novelty reference) and the relevance result (racc). This results in a single-generated reward incorporating both signal types. As per the previous limitation, Zhou teaches the relevance type component should be produced by an LLM configured to provide reference values. Therefore, by substituting Zhou’s LLM generated relevance result into Stamenković’s ‘stack(…)’ reward combination mechanism in place of or alongside racc arrives directly at this limitation, a single generated reward based on both the novelty reference candidate set and the LLM relevance result. This also matches Stamenković’s own disclosed technique of combining a novelty-based reward component with a relevance-based reward component into one value.)]; and
Zhang and Stamenković are analogous art because they are from the same field of endeavor of reinforcement-learning-based training of models that identify, select, or confirm target items relative to source items using retrieval- or ranking-quality metrics as the reward signal. They are further reasonably pertinent to the same problem of training such models using reinforcement learning because the target task cannot be trained with ordinary differentiable methods.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to substitute Stamenković’s multi-objective reward structure for Zhang’s single-objective NDCG reward within the existing state/action architecture.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”].
Zhang does not fully disclose, but Zhang with Zhou does teach:
receiving a relevance result that a relevance-reference model generates based on a prompt, the relevance result specifying whether the source item is relevant to the target item, the relevance-reference model being a model that serves as a reference for assessing relevance, the relevance-reference model being different than the plural-objective model [Zhou, pg. 12484, 4.2 Relevance Reward Model Training “we employ the large language model (LLM), LLaMA-13b (Touvron et al., 2023), to compute the generation likelihood of the document given a specific query.”. (note: this shows an LLM relevance oracle, distinct from the plural-objective model, receiving a query (document) prompt and returning a relevance result)];
Zhang and Zhou are analogous art because they are from the same field of endeavor of training models using reinforcement learning to select or evaluate items relative to a source item, using reward signals obtained from separate reference models. They are further reasonably pertinent to the same problem of generating a reliable training signal when direct labeled supervision for the objective being optimized is limited or unavailable.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to include Zhou’s LLM relevance signal into Zhang’s existing state and action framework.
The suggestion/motivation for doing so would have been Zhou’s relevance signal leads to improved performance [Zhou, pg. 12487, 6.2 Ablation Study on Relevance Annotators “the inclusion of the LLM annotator yields performance enhancements, indicating that leveraging the zero-shot capability of large language models can further augment the model’s ability to capture contextual dependencies.”].
Claim 18 is substantially similar in scope and spirit to claim 2. Therefore, the rejection of claim 2 is applied accordingly. Zhang further shows the method being implemented by an article of manufacture [Zhang, pg. 1, footnote "All our codes, data, and results are available at https://github.com/thunlp/ReInfoSelect." (note: Zhang's method exists as stored computer-readable instructions)]
Per claim 20, Zhang-Stamenković-Zhou discloses claim 17.
Zhang further teaches the selecting of the source item includes sampling the source item from a data store of source items [Zhang, pg. 4, 4 Experimental Methodology "About 100K anchors (from total 6 million collected) and their linked documents are randomly sampled as the weak supervision dataset." (note: Zhang discloses a data store (6 million collected anchors) from which source items (anchors) are randomly sampled)], and
Zhang does not fully disclose, but with Stamenković does teach wherein the selecting of the target item includes sampling the target item based on probability information produced by the plural-objective model based on the source item, the probability information describing likelihoods of different candidate items matching the source item. [Stamenković, pg. 3, 4 Model And Training "we cast the task of next item recommendation as a (self-supervised) multi-class classification problem and build a sequential model that receives user-item interaction sequence x1:t... and generates n classification logits yt+1 = [y1, y2, ..., yn] ∈ Rn"; pg. 4, 4 MODEL AND TRAINING “we train our recommendation model by optimizing the cross-entropy loss Ls based on the logits yt+1” (note: the quantity that is interpreted as this limitation is softmax(yt+1), the softmax normalized version of Stamenković’s classification logits. Stamenković states that these logits are used for a cross-entropy loss. The softmax step is a necessary, standard part of cross-entropy loss over multi-class logits.)]
Zhang, Stamenković, and Zhou are analogous art because they are from the same field of endeavor of training models using reinforcement learning to select or evaluate items relative to a source item, using reward signals obtained from separate reference models. They are further reasonably pertinent to the same problem of generating a reliable training signal when direct labeled supervision for the objective being optimized is limited or unavailable.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to include Stamenković’s multi-objective reward structure and Zhou’s LLM relevance signal into Zhang’s existing state and action framework.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”]. Incorporating Zhou’s relevance signal leads to improved performance [Zhou, pg. 12487, 6.2 Ablation Study on Relevance Annotators “the inclusion of the LLM annotator yields performance enhancements, indicating that leveraging the zero-shot capability of large language models can further augment the model’s ability to capture contextual dependencies.”].
Claim(s) 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zhang in view of Stamenković, in further view of Zhou and Karpukhin.
Per claim 19, Zhang-Stamenković-Zhou discloses claim 17.
Zhang does not fully disclose, but with Karpukhin does teach: the novelty-reference model has been trained using supervised training based on a training set [Karpukhin, pg. 6771, 3.2 Training “
PNG
media_image4.png
68
627
media_image4.png
Greyscale
”. (note: Karpukhin defines its training data D as a collection of instances, each comprising a source item and its associated and non-associated target items)] that specifies pairs of items that are considered associated [Karpukhin, pg. 6771, 3.2 Training “Each instance contains one question qi and one relevant (positive) passage
p
i
+
”. (note: the positive passage
p
i
+
is designated as relevant to the question qi within each training instance, which is pairs of items that are considered associated)] and pairs of items that are considered non-associated [Karpukhin, pg. 6771, 3.2 Training “…along with n irrelevant (negative) passages pi,j”. (note: the negative passages pi,j are designated irrelevant within each training instance, which is pairs of items that are considered non-associated)], based on a specified standard of association [Karpukhin, pg. 6771, 3.2 Training “passages relevant to a question may be given in a QA dataset, or can be found using the answer. All other passages in the collection, while not specified explicitly, can be viewed as irrelevant by default.”. (note: this shows the standard used for labeling passages positive versus negative by default, answer containment or relevance to a question. This is a specified standard of association.)].
Zhang, Stamenković, Zhou, and Karpukhin are analogous art because they are from the same field of endeavor of identifying or evaluating target items relative to a source item using learned representations or relevance signals. They are further reasonably pertinent to the same problem of generating a reliable training signal without relying solely on human-labeled supervision.
Before the effective filing date of the claimed invention, it would have been obvious to a
person of ordinary skill in the art to include Stamenković’s multi-objective reward structure, Zhou’s relevance signal, and Karpukhin’s indexed dual-encoder architecture into Zhang’s existing state and action framework.
The suggestion/motivation for doing so would have been to avoid the filter bubble problem that a purely relevance optimized reward produces, where users are repeatedly shown redundant, overly similar items, by also rewarding novelty and diversity in addition to relevance, which is explicitly stated by Stamenković [Stamenković, pg. 1, Abstract “…by focusing on item relevance, one pays a significant price in terms of other important metrics: users get stuck in a 'filter bubble'...”]. Incorporating Zhou’s relevance signal leads to improved performance [Zhou, pg. 12487, 6.2 Ablation Study on Relevance Annotators “the inclusion of the LLM annotator yields performance enhancements, indicating that leveraging the zero-shot capability of large language models can further augment the model’s ability to capture contextual dependencies.”]. Karpukhin’s indexed dual-encoder approach improves retrieval effectiveness. [Karpukhin, pg. 6770, 1 Introduction “we demonstrate that with the proper training setup, simply fine-tuning the question and passage encoders on existing question-passage pairs is sufficient to greatly outperform BM25.”.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Sayed M Shah whose telephone number is (571)272-9406. The examiner can normally be reached Monday-Friday 8:00 am - 4:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang can be reached at (571) 270-7092. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SAYED MUNEER SHAH/Examiner, Art Unit 2124
/MIRANDA M HUANG/Supervisory Patent Examiner, Art Unit 2124