DETAILED ACTION
This action is in response to the original filing on 05/23/2023. Claims 1-20 are pending and have been considered below.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 1 is objected to because of the following informalities: “training dataset” and “training set” are both used without any indication as to if these terms are equivalent or interchangeable. Further, “equal to an average” should be written as “equal to the average”. Appropriate correction is required.
Claim 4 is objected to because of the following informalities: The term “databases” appears twice when listing potential access points of the training set. Further, “cluster storages, cloud storages” should be rewritten to the less awkward form of “cluster/cloud storage.” Appropriate correction is required.
Claim 5 is objected to because of the following informalities: The term “corresponding embedding vectors” is a minor style inconsistency with the term as it appears in Claim 1, which reads as the singular “embedding vector”.
Claim 8 is objected to because of the following informalities: The term “displayed in response to the user changing the predetermined threshold value using the graphical user interface” is slightly awkward and could be simplified.
Claim 12 is objected to because of the following informalities: The term “dissimilarity score equals to” should read “dissimilarity score equal to”. Furthermore, the claim limitations make reference to “the training set” and never mentions a “training dataset”. As Claim 12 is a corresponding system to the method of Claim 1, appropriate correction should be made in one or both of these claims.
Claim 13 is objected to because of the following informalities: The term “databases” appears twice when listing potential access points of the training set. Further, “cluster storages, cloud storages” should be rewritten to the less awkward form of “cluster/cloud storage.” Appropriate correction is required.
Claim 14 is objected to because of the following informalities: The term “corresponding embedding vector” is a minor style inconsistency with the term as it appears in Claim 12, which makes reference to a possible plural context. Appropriate correction is required.
Claim 18 is objected to because of the following informalities: “prefemed” should recite “predefined”. Appropriate correction is required.
Claim 20 is objected to because of the following informalities: The term “higher or higher or equal” should be “higher than or equal to”. The claim is also very long and is missing parallelism in “configured to generate … and adding …”, which should be “configured to generate … and to add …” or “configured to generate …, and add …”.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: “centroid generator configured to”, “a dissimilarity score calculator configured to”, “outlier selector configured to”, “an embedding generator configured to” found in Claim 12, “sufficiency evaluator configured to” found in Claim 19, and “a data augmentation module configured to”, found in Claim 20.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding Claim 1, Claim 1 recites the limitation "gaining access to the training set for the neural network NN ". There is insufficient antecedent basis for this limitation in the claim. The preamble of claim 1 introduces "a training dataset," not "a training set," and the body of the claim uses both "training set" and "training dataset" interchangeably (e.g., "each element of the training dataset" vs. "each element of the training set"), rendering the claim indefinite as to whether these terms denote the same data collection. The claimed limitation is rejected on the basis of lacking antecedent basis.
For purposes of examination, "the training set" recited in the body of claim 1 is interpreted under BRI to refer to the same data collection introduced in the preamble as "a training dataset," because Spec [5]–[7] uses the terms "training set" and "training dataset" interchangeably to refer to the dataset that is processed for outlier identification.
Claim 1 further recites "a method for automatically identifying outliers in a training dataset for a neural network NN corresponding to a label." It is unclear whether the modifier "corresponding to a label" qualifies "a neural network NN" or "a training dataset." SPEC [15] uses different phrasing "a training set corresponding to a single label for a neural network" suggesting the modifier is intended to qualify the dataset. The ambiguity in claim 1 's grammar renders the metes and bounds of the claim unclear. For purposes of examination, “corresponding to a label” is interpreted under BRI as modifying “a training dataset,” i.e., the dataset comprises elements labeled with a single label, consistent with Spec [15]. The claimed limitation is rejected on the basis of ambiguous claim language.
Regarding Claim 2,
Claim 2 recites the limitation "the dissimilarity score for each element is calculated using a neural network trained using a metric learning method." Because parent claim 1 already introduces "a neural network NN," it is unclear whether the "neural network" of claim 2 refers to the same neural network NN or to a different, additional neural network. Spec [31]–[32] describes a metric-learning-trained neural network used to compute dissimilarity scores, but does not clearly distinguish whether this is the same network whose training set is being curated or a separate auxiliary network. The claimed limitation is rejected on the basis of ambiguous claim language.
For purposes of examination, "a neural network" recited in claim 2 is interpreted under BRI to refer to a neural network distinct from the "neural network NN" of claim 1, because Spec [31]–[32] describes a neural network trained using a metric learning method to compute dissimilarity scores between embeddings an auxiliary network used during outlier detection rather than the target network being trained on the cleaned dataset.
Claim 2 further recites “calculated using a neural network trained using a metric learning method”. This limitation is unclear because Claim 1 defines the dissimilarity score as a distance between the embedding vector corresponding to the element and the centroid; this claim may blur whether the neural network generates embedding vectors or calculates the score itself. The claimed limitation is rejected on the basis of ambiguous claim language.
For purposes of examination, the limitation recited is interpreted under BRI to mean the calculation rather than the generation of embeddings, because Spec [39] describes the generation of vectors to be undertaken by the embedded generator and the calculation to be undertaken by the score calculator.
Regarding Claims 4 and 13,
Claims 4 and 13 each recite a list of alternatives in which "databases" appears twice. Claim 4 recites "one or more user devices, databases, cluster storages, cloud storages, or databases." Claim 13 recites "one or more devices, databases, cluster storages, cloud storages, or databases." Each claim is indefinite because it cannot be determined whether the two recitations of "databases" refer to two distinct categories or are a drafting duplication. The Spec [6] and [16] contains the same duplication and provides no basis for distinguishing the two recitations.
For purposes of examination, the recitation of "databases" twice in the same alternative list is interpreted under BRI as a single recitation of "databases" alongside the other alternatives (user devices, cluster storages, cloud storages), because the SPEC [6] and [16] discloses the same alternative list and provides no basis for treating two different categories of databases.
Claim 4 further recites: “gaining access to one or more user devices, databases, cluster storages, cloud storages, or databases” This limitation is unclear whether gaining access to those components is the same as gaining access to the training set. The claimed limitation is rejected on the basis of ambiguous claim language.
For purposes of examination, the gaining access step is interpreted under BRI as gaining access to storages containing training data because the Spec [6] and [16] discloses the given list of storage capabilities, implying that is where the training data is held.
Regarding Claim 8, Claim 8 recites "in response to the user changing the predetermined threshold value" and claim 16 recites "an input device configured to allow the user to configure the predetermined threshold..." There is insufficient antecedent basis for "the user" in either claim. The dependency chain of claim 8 (claim 1 -> claim 7) and the dependency chain of claim 16 (claim 12) do not introduce "a user." Although claim 7 recites "a graphical user interface," no "user" is recited. There is insufficient antecedent basis for this limitation in the claim, and thus is rejected for this reason.
For purposes of examination, "the user" is interpreted under BRI to refer to a human operator of the system who interacts with the graphical user interface, as SPEC [42]–[43] describes the visual output device displaying information "to the user" and the input device allowing "the user to configure the predetermined threshold using a graphical user interface."
Regarding Claims 11, 18, and 20, Claims 11, 18, and 20 each recite "the predefined threshold." Claim 19 recites "the predetermined threshold." The independent claims (claim 1 and claim 12) introduce "a predetermined threshold value" not "a predefined threshold." The inconsistent use of "predefined" (claims 11, 18, 20) and "predetermined" (claims 1, 9, 12, 19) for what appears to be the same threshold renders these claims indefinite. Claim 19 compounds the ambiguity by reverting to "predetermined" in a claim that depends from claim 18 (which uses "predefined"); it is unclear whether the threshold of claim 19 is the same as the threshold of claim 18. There is insufficient antecedent basis for this limitation in the claim, and thus is rejected for this reason.
For purposes of examination, "the predefined threshold" is interpreted under BRI as referring to the same value as "a predetermined threshold value" introduced in claim 1/claim 12, because SPEC [40], [42]–[43] consistently uses one threshold value, configurable by user or dynamically generated, and uses "predetermined," "predefined," and "the threshold" interchangeably to refer to that single value.
Regarding Claim 12, Claim 12 recites "an embedding generator configured to generate an embedding vector for each element of the training set which is a numeric representation of that element." This limitation invokes 35 U.S.C. 112(f). The term "generator" is a generic placeholder/nonce term that does not by itself convey any structure to a person of ordinary skill in the art (Prong A); the term is modified by functional language ("configured to generate an embedding vector...") (Prong B); and the claim does not recite sufficient structure for performing the entire claimed function (Prong C).
However, the written description fails to disclose corresponding structure linked to the claimed function. Spec [37] describes the embedding generator only by repeating its function and does not disclose any algorithm, model architecture, training procedure, or other transformation method for converting an element into a numeric embedding vector. Spec [31] mentions metric learning to train a neural network for calculating dissimilarity scores, but this is not linked to the embedding-generation function. A general-purpose computer alone is insufficient corresponding structure for a computer-implemented function the specification must disclose an algorithm. The claim is therefore indefinite under 35 U.S.C. 112(b).
Applicant may: (a) amend the claim so that it does not invoke 35 U.S.C. 112(f); (b) amend the specification to expressly disclose the algorithm or other corresponding structure that performs the entire claimed function, without introducing new matter; or (c) amend the specification to clearly link the disclosed structure to the claimed function, without introducing new matter. The limitation is rejected on the basis of functional claiming with no corresponding structure 112(f).
For purposes of examination, "an embedding generator configured to generate an embedding vector for each element of the training set which is a numeric representation of that element" is interpreted under BRI to encompass any hardware or software component (including a general-purpose processor executing any algorithm or trained model) that produces a numeric vector representation of an input element.
Claim 12 further recites "a dissimilarity score calculator configured to calculate, for all the elements in the training set, a dissimilarity score, wherein the dissimilarity score equals to a distance between the embedding vector of the element and the centroid." This limitation invokes 35 U.S.C. 112(f) under the *Williamson* three-prong test for the same reasons as the embedding generator: "calculator" is a generic nonce term (Prong A), modified by functional language (Prong B), without sufficient recited structure (Prong C).
The written description's disclosure of corresponding structure is inadequate. Spec [39] states the score "equals to a distance between the embedding vector of the element and the centroid" and adds that "other functions equivalent to a distance...can also be utilized." "A distance" is not a single algorithm Euclidean distance, cosine distance, Manhattan distance, Mahalanobis distance, and other metrics are all candidates and the open-ended phrase "other functions equivalent to a distance" further leaves the algorithm unspecified. This is insufficient corresponding structure for the computer-implemented function, rendering the claim indefinite under 35 U.S.C. 112(b). The limitation is rejected on the basis of claiming functionality with no corresponding structure 112(f).
For purposes of examination, "a dissimilarity score calculator...wherein the dissimilarity score equals to a distance between the embedding vector of the element and the centroid" is interpreted under BRI to encompass any component that computes any distance metric (including Euclidean, cosine, Manhattan, Mahalanobis, or any other vector-space distance) between an element's embedding vector and the centroid.
Regarding Claim 14, Claim 14 recites "metadata indicating the relationship of each of the elements of the training set to the corresponding embedding vector." There is insufficient antecedent basis for "the relationship" because no prior recitation of "a relationship" appears in claim 14 or in claim 12 from which it depends. (Compare claim 5, which properly introduces this concept with the indefinite article: "metadata indicating a relationship of each of the elements..."). There is insufficient antecedent basis for this limitation in the claim, and thus is rejected for this reason.
For purposes of examination, "the relationship" is interpreted under BRI to refer to any associative correspondence between an element and its embedding vector (e.g., a pointer, identifier, offset, file name, or index linking the element to the vector), as the Spec [17] and [37] gives examples including "an offset of a video fragment corresponding to the embedding" and "the file name of the image corresponding to the embedding."
Regarding Claim 17, Claim 17 recites "the visual output device is further configured to refresh the list of the elements and the corresponding dissimilarity scores displayed to the user..." There is insufficient antecedent basis for both "the visual output device" and "the list." Claim 17 depends from claim 16, which depends from claim 12; neither claim 16 nor claim 12 recites a "visual output device" or a "list of...dissimilarity scores." The visual output device and list are introduced only in claim 15, a separate dependent claim from claim 12 that is not in claim 17's dependency chain. To establish proper antecedent basis, claim 17 should depend (directly or indirectly) from claim 15. There is insufficient antecedent basis for this limitation in the claim, and thus is rejected for this reason.
For purposes of examination, "the visual output device" recited in claim 17 is interpreted under BRI to refer to the visual output device introduced in claim 15, and "the list of the elements and the corresponding dissimilarity scores" is interpreted to refer to the "list of the dissimilarity scores with visual representations of the elements of the training set" recited in claim 15, as Spec [42]–[43] describes a single visual output device 212 performing both initial display and refresh.
Regarding Claim 19, Claim 19 recites "a sufficiency evaluator configured to identify if the training set is sufficient to train the neural network NN based on at least one sufficiency criterion after the elements with the dissimilarity score higher than or equal to the predetermined threshold are removed from the training set." This limitation invokes 35 U.S.C. 112(f). "Evaluator" is a generic nonce term (Prong A); the limitation is modified by functional language ("configured to identify if the training set is sufficient...") (Prong B); and the claim does not recite sufficient structure (Prong C).
Spec [44] discloses only that the criterion may be based on "the heuristic analysis of known systems of similar nature and similar number of degrees of freedom of NN," with a single example: "the heuristic rule may determine that the model requires N * D or D², where D is the number of degrees of freedom of the neural network, and N is a number." The broader concept of "heuristic analysis" is not defined; "N" is left as merely "a number" without specifying its value or how it is selected. A POSITA cannot determine, with reasonable certainty what algorithm the sufficiency evaluator must perform. The claim is therefore indefinite under 35 U.S.C. 112(b).
For purposes of examination, "a sufficiency evaluator configured to identify if the training set is sufficient to train the neural network NN based on at least one sufficiency criterion" is interpreted under BRI to encompass any component that applies any heuristic comparing training set size to a function of the number of degrees of freedom of the neural network, including but not limited to N×D or D² (where D is the number of degrees of freedom and N is any positive number), as described at Spec [44].
Regarding Claim 20, Claim 20 recites "the elements with the dissimilarity score higher or higher or equal to the predefined threshold are removed from the training set." The phrase "higher or higher or equal to" is grammatically nonsensical "higher" is repeated within a single comparison expression. It is unclear whether the intended comparison is (a) "higher than," (b) "higher than or equal to" (the comparison consistently used in claims 1, 9, 12, 18, and 19), or some other relation. The same drafting error appears in Spec [24] and [45], and so the specification does not resolve the ambiguity.
For purposes of examination, "higher or higher or equal to the predefined threshold" is interpreted under BRI as "higher than or equal to the predetermined threshold," consistent with the comparison expressly recited in independent claims 1 and 12 and dependent claims 9, 11, 18, and 19, which describe the same outlier-removal operation.
Regarding Claims 3, 5, 7, and 9-10, Claims 3, 5, 7, and 9-10 are also rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112(pre-AIA ), second paragraph, as being indefinite for depending on an indefinite parent claim.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Claims 1 and 12
Step 1: Claims 1 and 12 recite a method and a system; therefore, they are directed to the statutory categories of a method and a machine.
Step 2A Prong 1: The claims recite, inter alia: “generating, for each element of the training dataset, an embedding vector which is a numeric representation of the corresponding element; computing a centroid of all the embedding vectors of all the elements of the training set equal to an average of all the embedding vectors of all the elements of the training set; generating a dissimilarity score for each element of the training set by calculating a distance between the embedding vector corresponding to the element and the centroid; and marking the elements from the training set with embedding vectors having the dissimilarity score higher than or equal to a predetermined threshold value as outliers.” Under its broadest reasonable interpretation, these limitations encompass the mental process of observation or evaluation that is practically capable of being performed in the human mind with the aid of pen and paper.
Step 2A Prong 2: The judicial exception is not integrated into a practical application. The additional element of “gaining access to the training set for the neural network NN comprising a plurality of elements;” found in Claim 1 amounts to no more than well understood, routine, and conventional activity to apply an exception (MPEP 2106.05(d)), as “gaining access” to a training set is classed as WURC or as insignificant extra-solution activity (MPEP 2106.05(d)(II)(iv)(Storing and retrieving information in memory)). The additional element of “a data storage configured to store the training set;” found in Claim 12 amount to no more than generally linking the use of a judicial exception to a particular technological environment or field of use. (MPEP 2106.05(h)). The claim invokes computers or other machinery merely as a tool to perform an existing process.
Step 2B: The claim does not contain significantly more than the judicial exception. The additional element of “gaining access to the training set for the neural network NN comprising a plurality of elements;” found in Claim 1 amounts to no more than well understood, routine, and conventional activity to apply an exception (MPEP 2106.05(d)), as “gaining access” to a training set is classed as WURC or as insignificant extra-solution activity (MPEP 2106.05(d)(II)(iv)(Storing and retrieving information in memory)). The additional element of “a data storage configured to store the training set;” found in Claim 12 amount to no more than generally linking the use of a judicial exception to a particular technological environment or field of use. (MPEP 2106.05(h)). The claim invokes computers or other machinery merely as a tool to perform an existing process. Nothing in the claims provides significantly more than that abstract idea. As such, the claims are ineligible.
Claim 2
Step 1: Claim recites a method; therefore, it is directed to the statutory categories of a method.
Step 2A Prong 1: Claim 2 merely narrows the previously recited abstract limitations. For the reasons described above with respect to Claim 1, this judicial exception is not meaningfully integrated into a practical application, or significantly more than the abstract idea. The claims disclose similar limitations described for the independent claims above and do not provide anything more than the mental processes that are practically capable of being performed in the human mind with the assistance of pen and paper and mathematical concepts that are achievable through mathematical computation.
Step 2A Prong 2: Claim 2 recites the additional element of: “the dissimilarity score for each element is calculated using a neural network trained using a metric learning method.” amount to no more than mere instructions to implement an abstract idea or other exception on a computer. (see MPEP § 2106.05(f)).
Step 2B: The claims do not contain significantly more than the judicial exception.
Claim 3
Step 1: Claim recites a method; therefore, it is directed to the statutory categories of a method.
Step 2A Prong 1: The claim recites: “wherein the metric learning method implements at least one of a Center Loss, Triplet Loss, Contrastive Loss, Softmax Loss, A-Softmax Loss, Large Margin Cosine Loss (LMCL), or Arcface Loss methods.” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses mathematical processes that can be performed mentally, requiring specific mathematical calculations to perform the training.
Step 2A Prong 2 and Step 2B: There are no additional elements recited so the claim does not provide a practical application and is not considered to be significantly more. As such, the claims are patent ineligible.
Claim 4
Step 1: Claim recites a method; therefore, it is directed to the statutory categories of a method.
Step 2A Prong 1: Claim 4 merely narrows the previously recited abstract limitations. For the reasons described above with respect to Claim 1, this judicial exception is not meaningfully integrated into a practical application, or significantly more than the abstract idea. The claims disclose similar limitations described for the independent claims above and do not provide anything more than the mental processes that are practically capable of being performed in the human mind with the assistance of pen and paper and mathematical concepts that are achievable through mathematical computation.
Step 2A Prong 2: Claim 4 recites the additional element of: “wherein the gaining access to the training set for the neural network NN comprises gaining access to one or more user devices, databases, cluster storages, cloud storages, or databases.” amounts to no more than well understood, routine, and conventional activity to apply an exception (MPEP 2106.05(d)), as “gaining access” to a training set is classed as WURC or as insignificant extra-solution activity (MPEP 2106.05(d)(II)(iv)(Storing and retrieving information in memory)).
Step 2B: The claims do not contain significantly more than the judicial exception.
Claim 5
Step 1: Claim 5 recites a method; therefore, it is directed to the statutory categories of a method.
Step 2A Prong 1: The claim recites: “wherein generating the embedding vectors further comprises storing the embedding vectors, with metadata indicating a relationship of each of the elements of the training set to the corresponding embedding vectors.” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses mathematical processes that can be performed mentally, specifically organizing information and manipulating information through mathematical correlations (See MPEP 2106.04(a)(2)(A)(iv)) e.g. indicating a relationship of each of the elements of the training set to the corresponding embedding vectors.
Step 2A Prong 2 and Step 2B: There are no additional elements recited so the claim does not provide a practical application and is not considered to be significantly more. As such, the claims are patent ineligible.
Claim 6 and 15
Step 1: Claim recites a method and system; therefore, are directed to the statutory categories of a method and system.
Step 2A Prong 1: Claims 6 and 15 merely narrows the previously recited abstract limitations. For the reasons described above with respect to Claim 1 and 12, this judicial exception is not meaningfully integrated into a practical application, or significantly more than the abstract idea. The claims disclose similar limitations described for the independent claims above and do not provide anything more than the mental processes that are practically capable of being performed in the human mind with the assistance of pen and paper and mathematical concepts that are achievable through mathematical computation.
Step 2A Prong 2: Claim 6 recites the additional element of: “displaying a list of the dissimilarity scores with visual representations of the elements of the training set.” amounts to no more than additional elements recited at a high level of generality, e.g. no meaningful limitations of any particular method of displaying the recited information with any type of visual representations, thus represent insignificant extra solution activity of generic output, WURC similar to presenting offers and gathering statistics (see MPEP 2106.05(d)(II)). Claim 15 recites: “a visual output device configured to display a list of the dissimilarity scores with visual representations of the elements of the training set.” which amounts to no more than additional elements recited at a high level of generality, e.g. no meaningful limitations of any particular method of displaying the recited information with any type of visual representations, thus represent insignificant extra solution activity of generic output, WURC similar to presenting offers and gathering statistics (see MPEP 2106.05(d)(II)).
Step 2B: The claims do not contain significantly more than the judicial exception.
Claims 7 and 16
Step 1: Claim recites a method and system; therefore, are directed to the statutory categories of a method and system.
Step 2A Prong 1: Claims 7 and 16 merely narrows the previously recited abstract limitations. For the reasons described above with respect to Claim 1 and 12, this judicial exception is not meaningfully integrated into a practical application, or significantly more than the abstract idea. The claims disclose similar limitations described for the independent claims above and do not provide anything more than the mental processes that are practically capable of being performed in the human mind with the assistance of pen and paper and mathematical concepts that are achievable through mathematical computation.
Step 2A Prong 2: Claim 7 recites the additional element of: “wherein the predetermined threshold value is configurable by using a graphical user interface.” amounts to no more than mere instruction to apply the underlying mathematical concepts dealing with threshold value determination (See MPEP 2106.05(f)(2)). Claim 16 recites: “an input device configured to allow the user to configure the predetermined threshold using a graphical user interface.” amounts to no more than mere instruction to apply the underlying mathematical concepts dealing with threshold value determination (See MPEP 2106.05(f)(2)).
Step 2B: The claims do not contain significantly more than the judicial exception.
Claim 8 and 17
Step 1: Claim recites a method and system; therefore, are directed to the statutory categories of a method and system.
Step 2A Prong 1: Claims 8 and 17 merely narrows the previously recited abstract limitations. For the reasons described above with respect to Claim 7 and 16, this judicial exception is not meaningfully integrated into a practical application, or significantly more than the abstract idea. The claims disclose similar limitations described for the independent claims above and do not provide anything more than the mental processes that are practically capable of being performed in the human mind with the assistance of pen and paper and mathematical concepts that are achievable through mathematical computation.
Step 2A Prong 2: Claim 8 recites the additional element of: “refreshing a list of the elements and the corresponding dissimilarity scores displayed in response to the user changing the predetermined threshold value using the graphical user interface.” amount to no more than a recitation of the words "apply it" (or an equivalent) or are no more than mere instructions to implement an abstract idea or other exception on a computer since no particular inventive concepts are specified with high level recitation of “using the graphical user interface by user to change predetermined threshold value” (see MPEP § 2106.05(f)). Claim 17 recites: “wherein the visual output device is further configured to refresh the list of the elements and the corresponding dissimilarity scores displayed to the user in response to changing the predetermined threshold value.” amount to no more than a recitation of the words "apply it" (or an equivalent) or are no more than mere instructions to implement an abstract idea or other exception on a computer since no particular inventive concepts are specified with high level recitation of “using the graphical user interface by user to change predetermined threshold value” (see MPEP § 2106.05(f)).
Step 2B: The claims do not contain significantly more than the judicial exception.
Claim 9
Step 1: Claim 9 recites a method; therefore, it is directed to the statutory categories of a method.
Step 2A Prong 1: The claim recites: “wherein the marking the elements from the training set with the embedding vectors having the dissimilarity score higher than or equal to the predetermined threshold value further comprises removing marked elements from the training set.” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses mathematical processes that can be performed mentally, specifically organizing information and manipulating information through mathematical correlations (See MPEP 2106.04(a)(2)(A)(iv)) e.g. indicating a relationship of each of the elements of the training set to the corresponding embedding vectors and removing an identified element.
Step 2A Prong 2 and Step 2B: There are no additional elements recited so the claim does not provide a practical application and is not considered to be significantly more. As such, the claims are patent ineligible.
Claim 10 and 19
Step 1: Claims 10 and 19 recite a method and system; therefore, they are directed to the statutory categories of a method and a system.
Step 2A Prong 1: Claim 10 recites: “identifying if the training set with removed marked elements is sufficient to train the neural network NN based on at least one sufficiency criterion.” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses mathematical processes that can be performed mentally, specifically organizing information and manipulating information through mathematical correlations (See MPEP 2106.04(a)(2)(A)(iv)) e.g. indicating a relationship of each of the elements of the training set to the corresponding embedding vectors and determining if the elements are sufficient to train. Claim 19 recites: “a sufficiency evaluator configured to identify if the training set is sufficient to train the neural network NN based on at least one sufficiency criterion after the elements with the dissimilarity score higher than or equal to the predetermined threshold are removed from the training set.” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses mathematical processes that can be performed mentally, specifically organizing information and manipulating information through mathematical correlations (See MPEP 2106.04(a)(2)(A)(iv)) e.g. indicating a relationship of each of the elements of the training set to the corresponding embedding vectors and determining if the elements are sufficient to train.
Step 2A Prong 2 and Step 2B: There are no additional elements recited so the claim does not provide a practical application and is not considered to be significantly more. As such, the claims are patent ineligible.
Claims 11 and 20
Step 1: Claim recites a method and system; therefore, are directed to the statutory categories of a method and system.
Step 2A Prong 1: Claims 11 and 20 merely narrows the previously recited abstract limitations. For the reasons described above with respect to Claim 10 and 19, this judicial exception is not meaningfully integrated into a practical application, or significantly more than the abstract idea. The claims disclose similar limitations described for the independent claims above and do not provide anything more than the mental processes that are practically capable of being performed in the human mind with the assistance of pen and paper and mathematical concepts that are achievable through mathematical computation.
Step 2A Prong 2: Claim 11 recites the additional element of: “generating at least one additional element using a generative convolutional neural network trained on the remaining training set and adding the at least one additional element to the training set if the training set is determined to be insufficient to train the neural network NN after the elements with dissimilarity score higher than or equal to the predefined threshold have been removed from the training set.” amount to no more than mere instructions to apply an exception (see MPEP § 2106.05(f)(2)). Claim 20 recites: “a data augmentation module configured to generate at least one additional element using a generative convolutional neural network trained on the remaining training set and adding the at least one additional element to the training set if the training set is determined to be insufficient to train the neural network NN after the elements with the dissimilarity score higher or higher or equal to the predefined threshold are removed from the training set.” amount to no more than mere instructions to apply an exception (see MPEP § 2106.05(f)(2)).
Step 2B: The claims do not contain significantly more than the judicial exception.
Claim 13
Step 1: Claim recites a system; therefore, is directed to the statutory categories of a system.
Step 2A Prong 1: Claim 13 merely narrows the previously recited abstract limitations. For the reasons described above with respect to Claim 12, this judicial exception is not meaningfully integrated into a practical application, or significantly more than the abstract idea. The claims disclose similar limitations described for the independent claims above and do not provide anything more than the mental processes that are practically capable of being performed in the human mind with the assistance of pen and paper and mathematical concepts that are achievable through mathematical computation.
Step 2A Prong 2: Claim 13 recites the additional element of: “wherein the data storage comprises one or more devices, databases, cluster storages, cloud storages, or databases.” amount to no more than generally linking the use of a judicial exception to a particular technological environment or field of use (see MPEP § 2106.05(h)(iv)).
Step 2B: The claims do not contain significantly more than the judicial exception.
Claim 14
Step 1: Claim 14 recites a method; therefore, it is directed to the statutory categories of a method.
Step 2A Prong 1: The claim recites: “wherein the embedding generator is further configured to store the embedding vectors, with metadata indicating the relationship of each of the elements of the training set to the corresponding embedding vector.” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses mathematical processes that can be performed mentally, specifically organizing information and manipulating information through mathematical correlations (See MPEP 2106.04(a)(2)(A)(iv)) e.g. indicating a relationship of each of the elements of the training set to the corresponding embedding vectors.
Step 2A Prong 2 and Step 2B: There are no additional elements recited so the claim does not provide a practical application and is not considered to be significantly more. As such, the claims are patent ineligible.
Claim 18
Step 1: Claim 18 recites a method; therefore, it is directed to the statutory categories of a method.
Step 2A Prong 1: The claim recites: “remove the elements with dissimilarity score higher than or equal to the [predefined] threshold from the training set.” Under its broadest reasonable interpretation in light of the specification, this limitation encompasses mathematical processes that can be performed mentally, specifically organizing information and manipulating information through mathematical correlations (See MPEP 2106.04(a)(2)(A)(iv)) e.g. indicating a relationship of each of the elements of the training set to the corresponding embedding vectors and removing elements with a certain dissimilarity score.
Step 2A Prong 2: Claim 18 recites the additional element of: “wherein the outlier selector is further configured to” amount to no more than mere instructions to apply an exception (see MPEP § 2106.05(f)(2)).
Step 2B: The claim does not contain significantly more than the judicial exception. The additional element of “wherein the outlier selector is further configured to” amount to no more than mere instructions to apply an exception (see MPEP § 2106.05(f)(2)). The claim invokes computers or other machinery merely as a tool to implement an abstract idea/existing process. Nothing in the claims provides significantly more than that abstract idea. As such, the claims are ineligible.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 4-5, 9, 12-14, and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Pan et al. (US 20210083994 A1, hereinafter Pan) in view of Uva (US 20250053273 A1, hereinafter Uva).
Regarding Claim 1, Pan teaches a method for automatically identifying outliers in a training dataset for a neural network NN corresponding to a label, the method comprising:
generating, for each element of the training dataset, an embedding vector which is a numeric representation of the corresponding element;(Paragraph [0004] The training further includes generating training feature vectors from the training utterances. The training feature vectors include respective training feature vectors associated with each skill hot. The training further includes generating multiple set representations of the training feature vectors, where each set representation of the multiple set representations corresponds to a subset of the training feature vectors, and configuring the classifier model to compare input feature vectors to the multiple set representations.)
computing a centroid of all the embedding vectors of all the elements of the training set equal to an average of all the embedding vectors of all the elements of the training set; (Paragraph [0160] In some embodiments, block 1425 is the beginning of the iterative loop of the first round for determining clusters. Specifically, at block 1425, the training system 3509 may determine respective centroid locations (i.e., a respective location for each centroid) for the various clusters to be generated in this iteration; the quantity of centroid locations is equal to the count determined at block 1415. In some embodiments, in the first iteration of this loop, the training system 350 may choose a set of randomly selected centroid locations, having quantity equal to the count, within a feature space 630 into which the training utterances 615 fall) and
marking the elements from the training set with embedding vectors having the dissimilarity score higher than or equal to a predetermined threshold value as outliers. (Paragraph [0218] In some embodiments, the input utterance 303 may be deemed unrelated to any of the available skill bots 116 if the input feature vector 1710 is far enough away (e.g., exceeding a predetermined distance deemed to be excessive) from every bot vector. However, in additional or alternative embodiments, an input utterance 303 is not deemed unrelated based on dissimilarity to bot vectors alone. For instance, if the input feature vector 1710 is dissimilar (i.e., not sufficiently similar) to all the bot vectors, this is not necessarily a dispositive indication that the input utterance 303 belongs in the none class 316. (Items deemed too dissimilar are marked as belonging to a "none class".))
However, Pan fails to disclose: gaining access to the training set for the neural network NN comprising a plurality of elements; and generating a dissimilarity score for each element of the training set by calculating a distance between the embedding vector corresponding to the element and the centroid;
In the same field of endeavor, Uva teaches: gaining access to the training set for the neural network NN comprising a plurality of elements; (Paragraph [0083] In some embodiments, training sets may include a variety of subject matters, such as, as nonlimiting examples, medical report documents, electronic health records, entity documents, business documents, inventory documentation, emails, user communications, advertising documents, newspaper articles, and the like. In some embodiments, training sets of an LLM may include information from one or more public or private databases. As a non-limiting example, training sets may include databases associated with an entity. In some embodiments, training sets may include portions of documents associated with audiovisual data correlated to examples of outputs. (If a training set includes information from a private database, then at some point access must be gained to use anything in there due to its private nature.) and
generating a dissimilarity score for each element of the training set by calculating a distance between the embedding vector corresponding to the element and the centroid; (Paragraph [0117] With continued reference to FIG. 12, Computing device may be configured to generate a classifier using a K-nearest neighbors (KNN) algorithm. A “K-nearest neighbors algorithm” as used in this disclosure, includes a classification method that utilizes feature similarity to analyze how closely out-of-sample-features resemble training data to classify input data to one or more clusters and/or categories of features as represented in training data; this may be performed by representing both training data and input data in vector forms, and using one or more measures of vector similarity to identify classifications within training data, and to determine a classification of input data.)
It would have been obvious to a person having ordinary skill in the art before the effective filing date to have incorporated the concept of generating a dissimilarity score for each element in the set, as suggested by Uva into the system of Pan because both of these systems address the field of training data sets used by machine learning models. Doing so would improve Pan by allowing for a more efficient gathering or generating of training data used in machine learning models (Uva Paragraph [0004]).
Regarding Claim 4, the combination of Pan and Uva teaches the invention as claimed in Claim 1, including: wherein the gaining access to the training set for the neural network NN comprises gaining access to one or more user devices, databases, cluster storages, cloud storages, or databases. (Uva Paragraph [0083] In some embodiments, training sets may include a variety of subject matters, such as, as nonlimiting examples, medical report documents, electronic health records, entity documents, business documents, inventory documentation, emails, user communications, advertising documents, newspaper articles, and the like. In some embodiments, training sets of an LLM may include information from one or more public or private databases. As a non-limiting example, training sets may include databases associated with an entity. In some embodiments, training sets may include portions of documents associated with audiovisual data correlated to examples of outputs. (If a training set includes information from a private database, then at some point access must be gained to use anything in there due to its private nature.))
Regarding Claim 5, the combination of Pan and Uva teaches the invention as claimed in Claim 1, including: wherein generating the embedding vectors further comprises storing the embedding vectors, with metadata indicating a relationship of each of the elements of the training set to the corresponding embedding vectors (Pan paragraph [0004] In some embodiments, a system described herein includes a training system and a master bot. The training system is configured to train a classifier model. Training the classifier model includes accessing training utterances associated with skill bots, where the training utterances comprising respective training utterances associated with each skill bot of the skill bots. Each skill bot is configured to provide a dialog with a user. The training further includes generating training feature vectors from the training utterances. The training feature vectors include respective training feature vectors associated with each skill hot. Paragraph [0070] In certain embodiments, the master bot 114 is configured to be aware of the available skill bots 116. For instance, the master bot 114 may have access to metadata that identifies the various available skill bots 116, and for each skill bot 116, the capabilities of the skill bot 116 including the tasks that can be performed by the skill bot 116. (As each skill bot is defined by metadata and each skill bot is defined by its relationship to a vector, the metadata therefore indicates a relationship with the training vectors of the skill bots)).
Regarding Claim 9, the combination of Pan and Uva teaches the invention as claimed in Claim 1, including: wherein the marking the elements from the training set with the embedding vectors having the dissimilarity score higher than or equal to the predetermined threshold value further comprises removing marked elements from the training set. (Uva Paragraph [0121] Still referring to FIG. 12, computer, processor, and/or module may be configured to sanitize training data. “Sanitizing” training data, as used in this disclosure, is a process whereby training examples are removed that interfere with convergence of a machine-learning model and/or process to a useful result. For instance, and without limitation, a training example may include an input and/or output value that is an outlier from typically encountered values, such that a machine-learning algorithm using the training example will be adapted to an unlikely amount as an input and/or output; a value that is more than a threshold number of standard deviations away from an average, mean, or expected value, for instance, may be eliminated. Alternatively or additionally, one or more training examples may be identified as having poor quality data, where “poor quality” is defined as having a signal to noise ratio below a threshold value. Sanitizing may include steps such as removing duplicative or otherwise redundant data, interpolating missing data, correcting data errors, standardizing data, identifying outliers, and the like. In a nonlimiting example, sanitization may include utilizing algorithms for identifying duplicate entries or spell-check algorithms. (Removing data from a training set should it be found to be too dissimilar or erroneous from a configurable amount))
Regarding Claims 12-14 and 18, they are system claims that correspond to the method claims 1, 4-5, and 9 above. Therefore, they are rejected for the same reasons as method claims 1, 4-5, and 9 above.
Claims 2-3 are rejected under 35 U.S.C. 103 as being unpatentable over Pan in view of Uva, as applied in Claim 1 above, and in further view of BYDLON (EP 4154810 A1, hereinafter Bydlon).
Regarding Claim 2, the combination of Pan and Uva teaches the invention as claimed in Claim 1 above. The combination of Pan and Uva fails to disclose the dissimilarity score for each element is calculated using a neural network trained using a metric learning method.
In the same field of endeavor, Bydlon teaches: the dissimilarity score for each element is calculated using a neural network trained using a metric learning method ([0176] In embodiments, the auto-labeler ALL is operable to interface with the transformer stage TS' to obtain the latent representation of some or each frame or interval of frames (windows) s of the labelled data of the prototype set. For example, a trained encoder g.sub.e may be used, and the auto-labeler ALL computes a distance measure between some or all frames/windows within a known label, and some or all encoder generated representations of the unlabeled training data. One suitable distance measure may be the cosine similarity, or squared Euclidean distance (L2 norm), or other. (Network g calculates distance, which is synonymous with similarity. Network g is metric learning.))
It would have been obvious to a person having ordinary skill in the art before the effective filing date to have incorporated the concept of calculating a dissimilarity score using a neural network trained in a metric learning method as suggested by Bydlon into the combination of Pan and Uva because both of these system address the field of training data sets used by machine learning models. Doing so would improve the combination of Pan and Uva by including the advantages of efficient machine learning model training and applying it to a non-medical field (Bydlon Paragraph [0008]).
Regarding Claim 3, the combination of Pan, Uva, and Bydlon teaches the invention as claimed in Claim 2 including: wherein the metric learning method implements at least one of a Center Loss, Triplet Loss, Contrastive Loss, Softmax Loss, A-Softmax Loss, Large Margin Cosine Loss (LMCL), or Arcface Loss methods. (Bydlon Paragraph [0163] Network g may be trained in an iterative process, in which predicted (training output) labels ĉ.sub.t are compared to the ground-truth labels c.sub.i using a cost function F, such as cross-entropy loss, log loss, etc. Training continues in multiple iterations until convergence is achieved, i.e., predictions match the ground-truth (within a pre-define margin), where c.sub.i, ĉ.sub.t ∈ {1.. K}, and K denotes the number of classes. (Cross entropy loss falls under the same umbrella as softmax loss. For softmax loss you need a softmax activation layer and a cross entropy loss function both in operation. Therefore, it is part of the same structure))
Claims 6-8, 10-11, 15-17, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Pan in view of Uva, as applied in Claims 1 and 12 above, and in further view of KLAIMAN et al. (US 20210350176 A1, hereinafter Klaiman).
Regarding Claim 6, the combination of Pan and Uva teaches the invention as claimed in Claim 1 above. The combination of Pan and Uva fails to disclose the limitation of: displaying a list of the dissimilarity scores with visual representations of the elements of the training set.
In the same field of endeavor, Klaiman teaches: displaying a list of the dissimilarity scores with visual representations of the elements of the training set. (Paragraph [0261] To each of the tile pairs in the first set, e.g. pair 812, the label “similar” is assigned. To each of the tile pairs in the second set, e.g. pair 814, the label “dissimilar” is assigned. Paragraph [0051] Computing, for each of the images whose tiles were used for computing a respective similarity score, a respective similarity heat map as a function of the similarity scores, the pixel color and/or pixel intensities of the similarity heat map being indicative of the similarity of the tiles in the said image to the selected tile; and Paragraph [0052] displaying the similarity heat map.)
It would have been obvious to a person having ordinary skill in the art before the effective filing date to have incorporated the concept of calculating a dissimilarity score using a neural network trained in a metric learning method as suggested by Klaiman into the combination of Pan and Uva because both of these system address the field of training data sets used by machine learning models. Doing so would be beneficial and improve the combination of Pan and Uva, as easily viewing similar scores would allow professionals in a given field more efficient tools to practice their trade by specifying particular needs (Klaiman Paragraph [0005]).
Regarding Claim 7, the combination of Pan, Uva, and Klaiman teaches the invention as claimed in Claim 1, including: wherein the predetermined threshold value is configurable by using a graphical user interface. (Klaiman Paragraph [0034] According to embodiments, the GUI enables the user to select whether the relevance heat map is computed based on the numerical values of the MIL or based on the weights of the attention-MLL or based on the combined score. This may allow a user to identify if the output of the MIL and of the attention MLL in respect to the predictive power of a tile is significantly different. Paragraph [0097] According to other embodiments, the tile pairs are generated by selecting at least some or all of the tiles of each received image as a starting tile; for each starting tile, selecting all or a predefined number of “nearby tiles”, wherein a “nearby tile” is a tile within a first circle centered around the starting tile, whereby the radius of this circle is identical to a first spatial proximity threshold; for each starting tile, selecting all or a predefined number of “distant tiles”, wherein a “distant tile” is a tile outside of a second circle centered around the starting tile, whereby the radius of the said circle is identical to a second spatial proximity threshold; the selection of the predefined number can be performed by randomly choosing this number of tiles within the respective image area. The first and second proximity threshold may be identical, but preferably, the second proximity threshold is larger than the first proximity threshold. For example, the first proximity threshold can be 1 mm and the second proximity threshold can be 10 mm. Then, a first set of tile pairs is selected, whereby each tile pair comprises the start tile and a nearby tile located within the first circle. Each tile pair in the first set is assigned the label “similar” tissue patterns. In addition, a second set of tile pairs is selected, whereby each pair in the said set comprises the start tile and one of the “distant tiles”. Each tile pair in the second set is assigned the label “dissimilar” tissue patterns. For example, this embodiment may be used for creating “binary” labels “similar” or “dissimilar”. (The radius of the distance between tiles can be configured through usage of a GUI. These threshold values for dissimilarity are adjustable through usage of the GUI))
Regarding Claim 8, the combination of Pan, Uva, and Klaiman teaches the invention as claimed in Claim 7, including: refreshing a list of the elements and the corresponding dissimilarity scores displayed in response to the user changing the predetermined threshold value using the graphical user interface. (Klaiman Paragraph [0046] The feature that the tiles in the report gallery are selectable and a selection triggers the performing of a similarity search for identifying and displaying other tiles having a similar feature vector/tissue pattern as the user-selected tile may enable a user to freely select any image tile in the report tile gallery he or she is interested in... The display and the GUI can be refreshed automatically after the similarity search has completed. Paragraph [0047] According to some embodiments, the computation and display of the similarity search gallery comprises the computation and display of a similarity heat map. The heat map encodes similar tiles and respective feature vectors in colors and/or in pixel intensities. Image regions and tiles having similar feature vectors are represented in the heat map with similar colors and/or high or low pixel intensities. Hence, a user can quickly get an overview of the distribution of particular tissue pattern signatures in a whole slide image. The heat map can easily be refreshed simply by selecting a different tile, because the selection automatically induces a re-computation of the feature vector similarities based on the feature vector of the newly selected tile. (As established, the search for distant or near tiles requires selecting a proximity threshold. Therefore, as selecting a new tile automatically refreshes the computed similarity, the limitation is met.))
Regarding Claim 10, the combination of Pan, Uva, and Klaiman teaches the invention as claimed in Claim 9, including: identifying if the training set with removed marked elements is sufficient to train the neural network NN based on at least one sufficiency criterion. (Klaiman Paragraph [0145] In response to the selection of one or more high-predictive-power-tiles and/or artifact-tiles, automatically re-training the MIL program, thereby excluding the high-predictive-power-tiles and artifact-tiles from the training set. Paragraph [0146] These features may have the advantage that the re-trained MIL program may be more accurate, because the excluded artifact-tiles will not be considered any more during re-training. Hence, any bias in the learned transformation that was caused by tiles in the training data set depicting artifacts is avoided and removed by re-training the MIL program on a reduced version of the training data set that does not comprise the artifact-tiles. Paragraph [0147] Enabling a user to remove highly prognostic tiles from the training data set may be counter-intuitive but nevertheless provides important benefits: sometimes, the predictive power of some tissue patterns in respect to some labels is self-evident. (Identification and removal of particular tiles from a training set. The identification of if the remaining material is sufficient enough to train is automatically assumed to be true))
It would have been obvious to a person having ordinary skill in the art before the effective filing date to have incorporated the concept of identifying sufficiency of a dataset to train a neural network as suggested by Klaiman into the combination of Pan and Uva because both of these systems address the field of training data sets used by machine learning models. Doing so would be beneficial and improve the combination of Pan and Uva, as easily determining sufficiency of training datasets would allow professionals in a given field more efficient tools to practice their trade by specifying particular needs (Klaiman Paragraph [0005]).
Regarding Claim 11, the combination of Pan, Uva. And Klaiman teaches the invention as claimed in Claim 10, including: generating at least one additional element using a generative convolutional neural network trained on the remaining training set and adding the at least one additional element to the training set if the training set is determined to be insufficient to train the neural network NN after the elements with dissimilarity score higher than or equal to the predefined threshold have been removed from the training set. (Klaiman Paragraph [0175] A “fully convolutional neural network” as used herein is a neural network composed of convolutional layers without any fully-connected layers or multilayer perceptrons (MLPs) usually found at the end of the network. Paragraph [0206] According to one embodiment, the received images as well as the generated tiles are multi-channel images. The number of tiles can be increased for enriching the training data set by creating modified copies of existing tiles having different sizes, magnification levels, and/or comprising some simulated artifacts and noise. In some cases, multiple bags can be created by sampling the instances in the bag repeatedly as described herein for embodiments of the invention and placing the selected instances in additional bags. This “sampling” may also have the positive effect of enriching the training data set. (The enriching comprises creating copies of existing tiles and modifying them, such as by creating synthetic artifacts and noise. As stated in the analysis of Claim 10, the "sufficiency" step is automatically asumed to be true. The limitation is deemed to be met by the reference as the addition of synthetic data is up to the users discretion, therefore the user is the one deeming the data set as insuffcient.))
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Wu et al. (US 20220026519 A1) methods and systems for movement tracking, including vectors.
Reimann et al. (VN 10052622 B) predicting performance of animal feed via creation of a set of database vectors and removing quantified outliers from the set database.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUSTIN A CARDOSO whose telephone number is (571)272-8512. The examiner can normally be reached M-F 7:30 - 5:00, alternate Friday's off.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Welch can be reached at (571) 272-7212. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JUSTIN CARDOSO/
Patent Examiner, Art Unit 2143
/JENNIFER N WELCH/Supervisory Patent Examiner, Art Unit 2143