Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 7/27/2026 has been entered.
Response to Arguments
Applicant’s arguments with respect to claims 41-60 have been considered but are moot because the new ground of rejection does not rely on the same combination of references applied in the prior rejection of record for teachings or matters specifically challenged in the argument.
Drawings
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(5) because they include the following reference character(s) not mentioned in the description: 320 (Figure 3). Corrected drawing sheets in compliance with 37 CFR 1.121(d), or amendment to the specification to add the reference character(s) in the description in compliance with 37 CFR 1.121(b), are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either "Replacement Sheet" or "New Sheet" pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(d):
(d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph:
Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
Claims 42, 43, 51, and 52 are objected to under 35 U.S.C. 112(d) as failing to further limit the subject matter of the claim upon which they depend. Claim 42 recites "wherein the method further comprises sending a request for the at least one synthetic dataset," which is a limitation already recited in parent claim 41 ("sending a request for at least one synthetic dataset based on the at least one misclassification"). Claim 43 recites "wherein the at least one synthetic dataset is based on the misclassification," which is likewise already recited in parent claim 41. Claim 51 and claim 52 recite the corresponding limitations already present in parent claim 50. Appropriate correction is required, either by amending claims 42, 43, 51, and 52 to recite an additional limitation not already present in the claim from which they depend, or by cancellation.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 41-60 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Claims 41, 50 and 59 recite, in combination, the limitations: “assigning a classification score to each of the plurality of data types after the training of the at least one model on the at least one dataset”; “sending a request for at least one synthetic dataset based on the at least one misclassification”; “receiving the at least one synthetic dataset”; and “determining if the at least one misclassification is generated during the training of the at least one model on the at least one synthetic dataset based on the assigned classification score being below a particular threshold.” Claim 41 recites these as method steps, claim 50 as procedures performed by a computer arrangement executing stored instructions, and claim 59 as acts a processor is configured to perform. The written description analysis is the same for all three, and for the claims depending from them, because the recited subject matter is the same.
(1) The disclosure does not describe any embodiment that both assigns a classification score and requests and receives the synthetic dataset.
The disclosure divides the failure-feedback loop into three separate embodiments, each with its own figure:
Figure 6A (specification ¶[0059]): a dataset is received (602); a determination is made whether a misclassification is generated during training of the model on the dataset (604); “a classification score can be assigned to data types in the dataset after or during the training of the model” (606); a synthetic dataset is generated based on the misclassification (608); a determination is made whether the misclassification is still present during training on the synthetic dataset (610); and procedures 608 and 610 are repeated (612). This is the only figure in which a classification score appears — and in it the synthetic dataset is generated, not requested and received.
Figure 6B (specification ¶[0060]): a dataset including an identification of a plurality of data types is received (622); a misclassification of one of the data types is determined (624); a synthetic dataset is generated based on the misclassified data type (626); and a determination is made whether the misclassification is still present (628). No classification score, no threshold, no request, and no receipt.
Figure 6C (specification ¶[0061]): a dataset is received (632); a misclassification is determined (634); “a request for a synthetic dataset can be sent” (636); “the synthetic dataset can be received” (638); and a determination is made whether the misclassification is still present (640). This is the only figure in which a request is sent and a synthetic dataset is received — and it contains no classification score and no threshold.
The Summary of Exemplary Embodiments observes the same three-way division. The classification-score-and-threshold sentence at specification ¶[0009] — “[t]he misclassification(s) can be determined based on the assigned classification score being below a particular threshold” — is appended to the first summary embodiment of ¶[0008], whose act (c) is “generating a synthetic dataset(s) based on the misclassification(s)”, not requesting and receiving one. The second summary embodiment, at specification ¶[0011], likewise recites “generating a synthetic dataset(s)” and states only that “[a] classification score can be assigned to each of the data types” and that “[t]he misclassification(s) can be determined based on the assigned classification score” — with no threshold. The third summary embodiment, at specification ¶[0012], is the only passage in the Summary reciting “sending a request for a synthetic dataset(s) based on the misclassification” and “receiving the synthetic dataset(s)”; it recites no classification score, no threshold, and no requirement that the synthetic dataset include more of a particular data type.
Claims 41, 50 and 59 therefore take the request-and-receive framework from the ¶[0012] / Figure 6C embodiment and graft onto it the classification-score-below-a-threshold criterion of the ¶[0009] / Figure 6A embodiment. No passage of the specification and no figure describes that combination. Notably, the Detailed Description of the feedback loop — specification ¶[0028] through ¶[0039], the only portion of the disclosure that explains how the loop actually operates — does not mention a classification score at all, nor any threshold applied to a classification score. Picking features from separate disclosed embodiments and combining them in a manner the specification never describes does not convey possession of the resulting combination, and the disclosure provides no “blaze marks” directing one of ordinary skill to the claimed combination. See MPEP § 2163.02; Novozymes A/S v. DuPont Nutrition Biosciences APS, 723 F.3d 1336, 1349 (Fed. Cir. 2013); Purdue Pharma L.P. v. Faulding Inc., 230 F.3d 1320, 1326-27 (Fed. Cir. 2000); Ariad Pharmaceuticals, Inc. v. Eli Lilly & Co., 598 F.3d 1336, 1351 (Fed. Cir. 2010) (en banc).
(2) The disclosure does not describe determining the misclassification during training on the synthetic dataset based on a classification score assigned after the earlier training on the dataset.
The final limitation of claims 41, 50 and 59 requires that the determination made “during the training of the at least one model on the at least one synthetic dataset” be made “based on the assigned classification score being below a particular threshold.” By its antecedent, “the assigned classification score” is the score recited two limitations earlier as being assigned “after the training of the at least one model on the at least one dataset” — that is, a score produced by the first training pass, on the original dataset. The claims thus require the second-pass determination to be predicated on a first-pass score.
The specification describes the opposite arrangement. At specification ¶[0009], the score-below-threshold criterion is tied to determining the misclassification — the first determination, corresponding to procedures 604 and 606 of Figure 6A. The second determination in Figure 6A, procedure 610, is described as “a determination can be made as to whether the misclassification is still present during a training of the model on the synthetic dataset” (specification ¶[0059]), with no reference whatsoever to a classification score or a threshold. Nothing in the disclosure describes storing, carrying forward, or re-using a score assigned after the first training pass as the criterion governing the second-pass determination. Where the Detailed Description addresses the second-pass determination, it describes a fresh inquiry into causation rather than a threshold comparison against a previously-assigned score: “[i]f a misclassification is still present, then a determination can again be made as to what data caused the misclassification” (specification ¶[0035]).
(3) The failure-detection criteria that the specification does describe are different from the claimed criterion. Where the Detailed Description explains how a misclassification or failure is actually detected, it discloses mechanisms other than a classification score compared against a threshold: a count of misclassifications: “during the training of the model, a count of the number of misclassifications can be determined (e.g., continuously, or at predetermined intervals). Once the count reaches a certain threshold number, the misclassified data can be used to separately generate more samples of the same type of data” (specification ¶[0031]);
a statistical significance test: “[t]he result can be statistically significant, by the standards of the study, when p < α” (specification ¶[0032]); for non-categorical datasets, a user-defined function measuring departure from an expected value, including “the variance, which can be based on the number of standard deviations the produced value is away from the expected value (e.g., one, two, or three standard deviations)” (specification ¶[0030]); and direct identification by the training model, or verification by a separately-trained verification model (specification ¶[0029]).
None of these is a classification score assigned to each of a plurality of data types and compared against a threshold. That the specification describes these alternative and differently-structured criteria in operational detail, while the claimed criterion appears nowhere outside the bare labels at ¶[0009], ¶[0011] and procedure 606 of ¶[0059], reinforces that the inventors were not in possession of the claimed score-and-threshold-gated request-and-receive loop at the time of filing.
(4) The specification does not describe what the classification score or the threshold is. Beyond naming the terms, the specification does not describe how a classification score is computed or on what scale; how a single score is assigned to a data type — a category — as distinct from an individual data sample or an individual prediction; what the “particular threshold” is, what values it may take, or how it is selected; or any example of the score or the threshold in operation. The specification contains no working example and no experimental data of any kind. For a computer-implemented claim, a bare functional recitation of a result to be achieved, without disclosure of the algorithm that achieves it, does not demonstrate possession. See MPEP § 2163.02; Vasudevan Software, Inc. v. MicroStrategy, Inc., 782 F.3d 671, 681-82 (Fed. Cir. 2015).
Applicant is reminded that although claims as originally filed constitute their own written description, that presumption is rebuttable where the claimed scope exceeds what the specification as a whole reasonably conveys. See MPEP § 2163(II)(A)(3)(a)(i); In re Koller, 613 F.2d 819, 823 (CCPA 1980). The recitations at ¶[0009], ¶[0011] and ¶[0012] are restatements of claim language within the Summary of Exemplary Embodiments rather than substantive description, and the Detailed Description does not supply the missing description of the claimed combination.
Claims 42-49, 51-58 and 60 are rejected under the written description requirement based on their dependency from the rejected independent claims. None of the dependent claims supplies the missing description. In particular, claims 42 and 51 (“sending a request for the at least one synthetic dataset”) and claims 43 and 52 (“the at least one synthetic dataset is based on the misclassification”) merely restate limitations already recited in the independent claims from which they depend, and claims 49, 58 and 60 (“the at least one model is a machine learning procedure”) narrow the model without addressing the undescribed combination.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 41-60 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. The specific indefiniteness issues, and the claims to which they apply, are set forth below.
Claims 41, 50, and 59 recite "training [or: train] the at least one model on a combination of the at least one dataset and the synthetic dataset." There is insufficient antecedent basis for "the synthetic dataset" in these claims, as each claim previously introduces only "at least one synthetic dataset" (not "a synthetic dataset"). It is unclear whether "the synthetic dataset" refers to the same entity as "the at least one synthetic dataset" recited earlier in the claim, or to a different, unclaimed synthetic dataset. This deficiency is carried into dependent claims 42-49 (depending from claim 41), 51-58 (depending from claim 50), and 60 (depending from claim 59), none of which
For purposes of examination, "the synthetic dataset" is interpreted under BRI to refer to the same entity as "the at least one synthetic dataset" received earlier in each claim, as the specification at ¶[0033]-[0035] describes only a single synthetic dataset being combined with the initial dataset for retraining at this stage of the process.
Claims 41, 50, and 59 recite determining a misclassification "based on the assigned classification score being below a particular threshold." The term "particular threshold" is a relative/undefined term. It is not defined by the claims, and the specification does not provide an objective standard for ascertaining the requisite threshold value — at most, ¶[0031] of the specification refers generically to a "threshold number" without further definition. One of ordinary skill in the art would not be reasonably apprised of the scope of the invention. This deficiency is carried into dependent claims 42-49, 51-58, and 60. For purposes of examination, "a particular threshold" is interpreted under BRI to encompass any pre-established or dynamically-set numerical or qualitative boundary value used for comparison against the classification score, consistent with the specification's generic reference at ¶[0031] to a "certain threshold number."
Claims 41, 50, and 59 each recite, as the penultimate limitation, training the model "on a combination of the at least one dataset and the synthetic dataset," and then, as the final limitation, determining whether a misclassification is generated "during the training of the at least one model on the at least one synthetic dataset" — omitting reference to "the at least one dataset" or the "combination" used in the immediately preceding limitation. It is unclear whether the final limitation refers back to the same training operation recited in the preceding limitation, or to a separate, unclaimed training operation involving the synthetic dataset alone. This ambiguity prevents a person of ordinary skill in the art from ascertaining the metes and bounds of the final limitation with reasonable certainty.
This deficiency is carried into dependent claims 42-49, 51-58, and 60. For purposes of examination, the phrase "the training of the at least one model on the at least one synthetic dataset" in the final limitation of each independent claim is interpreted under BRI as referring to the same training operation recited in the immediately preceding limitation — i.e., training on the combination of the at least one dataset and the synthetic dataset — as the specification at ¶[0035] describes retraining the training model "on this combined dataset."
Claims 41, 50, and 59 each recite assigning a classification score "after the training of the at least one model on the at least one dataset" — i.e., after only the first (initial) training operation — but then rely on "the assigned classification score" to determine whether a misclassification is generated during the second training operation (training on the combination of the dataset and synthetic dataset). No limitation recites re-assigning or updating the classification score after the second training operation. It is unclear how a classification score computed only after the first training operation forms the basis for assessing the outcome of a subsequent, second training operation, rendering the scope of the final limitation indefinite.
This deficiency is carried into dependent claims 42-49, 51-58, and 60. For purposes of examination, "the assigned classification score" as used in the final limitation of each independent claim is interpreted under BRI to encompass any classification score resulting from either the first or second training operation, consistent with the specification's disclosure at ¶[0029]-[0031] of classification/misclassification assessment as an ongoing, iterative process performed throughout training.
Claims 43 and 52 recite "wherein the at least one synthetic dataset is based on the misclassification." Parent claims 41 and 50, respectively, recite only "at least one misclassification" (not "a misclassification"). There is insufficient antecedent basis for "the misclassification" in claims 43 and 52. For purposes of examination, "the misclassification" in claims 43 and 52 is interpreted under BRI to refer to "the at least one misclassification" recited in parent claims 41 and 50, respectively.
Claims 44 and 53 recite "sending a request for additional data related to a particular one of the data types." Parent claims 41 and 50, respectively, recite only "the plurality of data types," not "the data types" as a standalone noun phrase, creating ambiguity as to whether "the data types" refers to the same plurality of data types recited in the parent claim or to some other, unclaimed set of data types. For purposes of examination, "the data types" in claims 44 and 53 is interpreted under BRI to refer to "the plurality of data types" recited in parent claims 41 and 50, respectively.
Claims 46 and 55 recite "wherein the at least one model is trained on at least one non-synthetic dataset and at least one further synthetic dataset." These terms are newly introduced without antecedent basis tying them to the previously recited "at least one dataset" or "at least one synthetic dataset" of parent claims 41 and 50, respectively. It is unclear whether these newly recited datasets are the same as, or in addition to, the datasets already recited in the parent claim, rendering the scope of claims 46 and 55 indefinite. For purposes of examination, "at least one non-synthetic dataset" is interpreted under BRI to refer to the same "at least one dataset" of the respective parent claim to the extent it consists of real (non-synthetic) data, and "at least one further synthetic dataset" is interpreted under BRI to refer to an additional synthetic dataset generated beyond "the at least one synthetic dataset" of the parent claim, consistent with the specification's disclosure at ¶[0038] of generating further/additional synthetic datasets when misclassifications persist.
Appropriate correction is required.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 41-60 are rejected under 35 USC 103 as being unpatentable over US Pat. Pub. No. 2018/0268255 A1 to Surazhsky et al. (hereinafter Surazhsky) in view of SMOTE: Synthetic Minority Over-sampling Technique to Chawla et al. (previously cited, hereinafter Chawla).
Per claim 41, Surazhsky discloses A method performed by a computer hardware arrangement (Surazhsky: ¶[0051]-[0052], FIG. 5:500…Surazhsky carries out the disclosed training method on a computing system 500 assembled from a processor 510, a memory 520, a storage device 530 and input/output devices 540 interconnected by a system bus 550, and states that computing system 500 implements the training engine 110 and the neural network engine 140 that perform every operation of the method, which constitutes the computer hardware arrangement performing the method under BRI, “the computing system 500 can include a processor 510, a memory 520, a storage device 530, and input/output devices 540”), the method comprising:
training at least one model on at least one dataset including a plurality of data types (Surazhsky: ¶[0039]…Surazhsky trains a convolutional neural network by processing a mixed training set built from a non-synthetic image and the synthetic images derived from it, which constitutes training at least one model on at least one dataset under BRI, “the training controller 218 may train the convolutional neural network by at least processing, with the convolutional neural network, a mixed training set that includes the non-synthetic image 300, the first synthetic image 330, the second synthetic image 332, the third synthetic image 334, and the fourth synthetic image 336”; ¶[0037]…every image in that mixed training set carries a category label such as vehicle or sign, and the collection of such labelled categories within the one dataset constitutes the plurality of data types under BRI, “the first synthetic image 330, the second synthetic image 332, the third synthetic image 334, and the fourth synthetic image 336 should have the same labels (e.g., vehicle, sign) as the non-synthetic image 300”);
determining at least one misclassification of one of the plurality of data types (Surazhsky: ¶[0040]…Surazhsky's performance auditor 214 inspects the result of the training pass and identifies that the trained network misclassifies the images belonging to one particular category of the training set - those subjected to a changed perspective - which constitutes determining at least one misclassification of one of the plurality of data types under BRI, “the performance auditor 214 may determine, based on a result of the processing of a mixed training set performed by the convolutional neural network, that the convolutional neural network misclassifies synthetic images from the mixed training set that have been subject to certain modifications”; ¶[0040]…the auditor's finding is category-specific rather than dataset-wide, isolating the single data type the model fails on, “Accordingly, the performance auditor 214 may determine that the convolutional neural network is unable to successfully classify synthetic images having changed perspectives”);
…
sending a request for at least one synthetic dataset based on the at least one misclassification; receiving the at least one synthetic dataset (Surazhsky: ¶[0026] …Surazhsky deploys the components that audit the model and the components that manufacture the synthetic data on separate remote platforms that reach one another only as a client and a server over network 120; between two machines joined solely by a network there is no mechanism by which one obtains data held by the other except by transmitting a demand for it and taking delivery of the response, so the act of sending a request for the synthetic data and the act of receiving that data are necessarily performed, not merely made possible, by Surazhsky's architecture, “the functionalities of the training engine 110 and/or the neural network engine 140 may be accessed (e.g., by the client device 130) as a remote service (e.g., a cloud application) via the network 120”; ¶[0056]…Surazhsky characterizes the interaction between its remote components in client-server terms, which is a request-and-response exchange by definition, “The programmable system or computing system may include clients and servers. A client and server are generally remote from each other and typically interact through a communication network”; ¶[0047]…the demand for synthetic data issues only upon, and is populated by, the performance auditor's misclassification finding, so the request that necessarily accompanies that transaction is one based on the at least one misclassification, “the training engine 110 (e.g., the synthetic image generator 210) may generate additional synthetic images having changed perspectives, when the training engine 110 (e.g., the performance auditor 214) determines that the convolutional neural network is unable to correctly classify synthetic images having changed perspectives”; ¶[0021]…Surazhsky further establishes that the components are joined by a network rather than by in-process calls, which is the structural premise of the request-and-response inference, “a training engine 110 may be communicatively coupled, via a wired and/or wireless network 120, to client device 130 and/or a neural network engine 140”);
wherein the at least one synthetic dataset includes more of a particular one of the plurality of data types than the at least one dataset, wherein the particular one of the plurality of data types is based on the at least one misclassification (Surazhsky: ¶[0047]…the synthetic data Surazhsky manufactures after the audit consists of additional images of the one category the network failed on, so that the newly generated set holds more of that category than the mixed training set that preceded it, which constitutes a synthetic dataset containing more of a particular data type than the at least one dataset under BRI, “the training engine 110 (e.g., the synthetic image generator 210) may generate additional synthetic images having changed perspectives, when the training engine 110 (e.g., the performance auditor 214) determines that the convolutional neural network is unable to correctly classify synthetic images having changed perspectives”; ¶[0050]…Surazhsky selects the category to be replenished from the misclassification finding itself, which constitutes the particular data type being based on the at least one misclassification, “the training engine 110 (e.g., the synthetic image generator 210) may generate additional training data, which may include synthetic images having the modifications that are applied to the synthetic images misclassified by the convolutional neural network”);
training the at least one model on a combination of the at least one dataset and the synthetic dataset (Surazhsky: ¶[0018]…Surazhsky's training corpus is by construction a union of the non-synthetic images and the synthetic images generated from them, which constitutes a combination of the at least one dataset and the synthetic dataset under BRI, “a convolutional neural network may be trained using a mixed training set that includes both synthetic and non-synthetic images”; ¶[0050]…training does not restart on the new synthetic data alone but continues on the already-trained network, so the model is trained on the original images together with the newly generated ones, “the training engine 110 (e.g., the training controller 212) may continue to train the convolutional neural network by at least processing the additional training data with the convolutional neural network”); and
determining if the at least one misclassification is generated during the training of the at least one model on the at least one synthetic dataset… (Surazhsky: ¶[0049]…after the network has been trained on the newly generated synthetic images Surazhsky re-audits the result of processing those very images to see whether the failure persists, which constitutes determining if the misclassification is generated during the training of the model on the synthetic dataset under BRI, “The training engine 110 may determine, based at least on a result of processing the one or more additional synthetic images, whether a performance of the machine learning model meets a threshold value”; ¶[0050]…a persisting failure returns the process to the generation step so the determination governs whether the loop repeats, “if the training engine 110 determines that the performance of the machine learning model does not meet the threshold value (413-N), the process 400 may continue at operation 410”).
Surazhsky does not expressly disclose, but Chawla does teach:
assigning a classification score to each of the plurality of data types after the training of the at least one model on the at least one dataset (Chawla: p. 345, Section 5.5…Chawla evaluates the trained classifier by computing and reporting a separate accuracy figure for every category in the dataset rather than a single aggregate figure, and such a per-category figure computed once training is complete constitutes a classification score assigned to each of the plurality of data types after the training under BRI, “Acc+ is the accuracy on positive (minority) examples and Acc− is the accuracy on the negative (majority) examples”);
…based on the assigned classification score being below a particular threshold (Chawla: p. 345, Section 5.4…Chawla judges the trained classifier's per-class score against an explicitly settable threshold and sweeps that threshold across a series of values, so that whether a category is treated as correctly or incorrectly classified turns on whether its score falls below the threshold in force, which constitutes determining the misclassification based on the assigned classification score being below a particular threshold under BRI, “We experimented by setting the decision thresholds at the leaves for the C4.5 decision tree learner at 0.5, 0.45, 0.42, 0.4, 0.35, 0.32, 0.3, 0.27, 0.25, 0.22, 0.2, 0.17, 0.15, 0.12, 0.1, 0.05, 0.0”; p. 345, Section 5.4…Chawla states the threshold comparison in terms of the class score at the leaf, “a leaf could classify examples as the minority class even if more than 50% of the training examples at the leaf represent the majority class”).
Surazhsky and Chawla are analogous art because they are from the same field of endeavor, specifically the supervised training of machine-learning classifiers on datasets in which one or more categories of data are inadequately represented. They address the same problem of curing a trained classifier's failure on a deficient category by manufacturing additional artificial examples of that category and retraining on the enlarged dataset.
Before the effective filing date of the claimed invention, it would have been obvious to a PHOSITA to score each data type of Surazhsky's mixed training set once training is complete and to condition Surazhsky's generation of additional synthetic data on that per-type score falling below a threshold, in the manner Chawla teaches, as claim 41 recites.
The suggestion/motivation for doing so would have been provided by Chawla itself, which teaches that a classifier built on a dataset whose categories are not approximately equally represented performs measurably worse on the under-represented category and that the remedy is to enlarge that category with synthetic examples, “An approach to the construction of classifiers from imbalanced datasets is described. A dataset is imbalanced if the classification categories are not approximately equally represented” (Chawla: p. 321, Abstract), and which supplies the per-category measurement by which the deficient category is identified, “The minority class is over-sampled by taking each minority class sample and introducing synthetic examples along the line segments joining any/all of the k minority class nearest neighbors” (Chawla: p. 328, Section 4.2). Surazhsky supplies the reciprocal incentive, teaching that aiming the additional synthetic data at the category the model actually fails on shortens training, “It should be appreciated that training the convolutional neural network in this targeted manner may accelerate the training process” (Surazhsky: ¶[0020]). Furthermore, this is the application of a known technique to a known method ready for improvement to yield predictable results, the rationale of MPEP § 2143(D).
Per claim 42, Surazhsky combined with Chawla discloses claim 41. Surazhsky further teaches the method further comprises sending a request for the at least one synthetic dataset (Surazhsky: ¶[0026], ¶[0047]…Surazhsky's synthetic image generator resides on a remote platform reached over network 120 and produces the additional data only when the performance auditor calls for it, so obtaining that data necessarily proceeds by a transmitted request as addressed in the intrinsic-disclosure paragraph above, “the functionalities of the training engine 110 and/or the neural network engine 140 may be accessed (e.g., by the client device 130) as a remote service (e.g., a cloud application) via the network 120”).
Per claim 43, Surazhsky combined with Chawla discloses claim 41. Surazhsky further teaches the at least one synthetic dataset is based on the misclassification (Surazhsky: ¶[0050]…Surazhsky selects the modification category to be reproduced in the additional training data from the very images the network misclassified, which constitutes the synthetic dataset being based on the misclassification under BRI, “the training engine 110 (e.g., the synthetic image generator 210) may generate additional training data, which may include synthetic images having the modifications that are applied to the synthetic images misclassified by the convolutional neural network”).
Per claim 44, Surazhsky combined with Chawla discloses claim 41. Surazhsky further teaches the method further comprises sending a request for additional data related to a particular one of the data types (Surazhsky: ¶[0047]…The data Surazhsky calls for after the audit is confined to one nominated category - images having changed perspectives - so the demand transmitted to the generator is one for additional data related to a particular one of the data types under BRI, “the training engine 110 (e.g., the synthetic image generator 210) may generate additional synthetic images having changed perspectives, when the training engine 110 (e.g., the performance auditor 214) determines that the convolutional neural network is unable to correctly classify synthetic images having changed perspectives”).
Per claim 45, Surazhsky combined with Chawla discloses claim 41. Surazhsky further teaches the at least one dataset includes one of (i) only real data, (ii) only synthetic data, or (iii) a combination of real data and synthetic data (Surazhsky: ¶[0018]…Surazhsky's mixed training set holds both non-synthetic images and synthetic images, which satisfies alternative (iii) of this Markush-style recitation under BRI, “a convolutional neural network may be trained using a mixed training set that includes both synthetic and non-synthetic images”).
Per claim 46, Surazhsky combined with Chawla discloses claim 41. Surazhsky further teaches the at least one model is trained on at least one non-synthetic dataset and at least one further synthetic dataset (Surazhsky: ¶[0018]…The same mixed training set supplies the network with a body of non-synthetic images and a body of synthetic images, which constitutes training on at least one non-synthetic dataset and at least one further synthetic dataset under BRI, “a convolutional neural network may be trained using a mixed training set that includes both synthetic and non-synthetic images”).
Per claim 47, Surazhsky combined with Chawla discloses claim 41. Surazhsky further teaches the at least one dataset includes an identification of each of the plurality of data types in the at least one dataset (Surazhsky: ¶[0018]…Every image in Surazhsky's training set carries a classification label naming the category it belongs to, and the synthetic images inherit those labels, which constitutes the dataset including an identification of each of the plurality of data types under BRI, “The non-synthetic images may be labeled with classifications that correspond to the objects depicted in these images”).
Per claim 48, Surazhsky combined with Chawla discloses claim 41. Surazhsky further teaches the method further comprises verifying an accuracy of the at least one model using at least one verification model (Surazhsky: ¶[0049]…Surazhsky's performance auditor 214 checks the trained network's accuracy against the error or cost function associated with that network, and that error-function model applied by a component separate from the network under test constitutes at least one verification model under BRI, “the training engine 110 (e.g., the performance auditor 214) may further gauge the performance of the convolutional neural network based on the error function or cost function associated with the convolutional neural network”; ¶[0017]…Surazhsky further discloses a held-out validation set processed separately from the training data, which supplies the independent yardstick against which the auditor verifies accuracy, “validation data may be similarly labeled data that is not used as training data”).
Per claim 49, Surazhsky combined with Chawla discloses claim 41. Surazhsky further teaches the at least one model is a machine learning procedure (Surazhsky: ¶[0022]…Surazhsky's model is a convolutional neural network, which is a machine learning procedure under BRI, “the neural network engine 140 may be configured to implement one or more machine learning models including, for example, a convolutional neural network”).
Claims 50-58 are substantially similar in scope and spirit to claims 41-49. Therefore, the rejections of claims 41-49 are applied accordingly. Suranzhsky further discloses A non-transitory computer-accessible medium having stored thereon computer-executable instructions, wherein, when a computer arrangement executes the instructions, the computer arrangement is configured to perform procedures… (Surazhsky: ¶[0057]…Surazhsky stores the program code that carries out the disclosed training operations on a machine-readable medium and states that the medium holds those instructions non-transitorily on solid-state memory or a magnetic hard drive, which constitutes the non-transitory computer-accessible medium having stored thereon computer-executable instructions under BRI, “The machine-readable medium can store such machine instructions non-transitorily, such as for example as would a non-transient solid-state memory or a magnetic hard drive or any equivalent storage medium”; ¶[0052], FIG. 5:510…the processor 510 of computing system 500 executes those stored instructions to perform the operations, which constitutes the computer arrangement being configured by the instructions to perform the recited procedures, “The processor 510 is capable of processing instructions stored in the memory 520 and/or on the storage device 530 to display graphical information for a user interface provided via the input/output device 540”).
Claims 59 and 60 are substantially similar in scope and spirit to claims 41 and 49, respectively. Therefore, the rejections of claims 41 and 49 are applied accordingly.
Surazhsky further discloses A system comprising (Surazhsky: ¶[0051], FIG. 5:500…Surazhsky's computing system 500 realizes the training engine 110 and the neural network engine 140 and all components therein, which constitutes the recited system under BRI, “the computing system 500 can be used to implement the training engine 110, the client device 130, the neural network engine 140, and/or any components therein”): a processor configured to … (Surazhsky: ¶[0052], FIG. 5:510…the processor 510 of that computing system executes the instructions that carry out each recited operation, which constitutes a processor configured to perform them under BRI, “The processor 510 is capable of processing instructions for execution within the computing system 500”).
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claims 41-60 are rejected on the ground of nonstatutory double patenting as being unpatentable over claims 1-5, 8-9 and 15-20 of U.S. Patent No. 11,631,032. Although the claims at issue are not identical, they are not patentably distinct from each other because the claims of the instant application and the claims of U.S. Patent No. 11,631,032 are directed to the same failure-feedback training scheme, and the instant claims merely recite a broadened, obvious variation of the patented claims. Instant claims 41, 50 and 59 recite training a model on a dataset having a plurality of data types, determining a misclassification of one of the data types, assigning a classification score to each data type after the training, sending a request for a synthetic dataset based on the misclassification, receiving the synthetic dataset, the synthetic dataset including more of the particular data type on which the misclassification was based, retraining the model on a combination of the dataset and the synthetic dataset, and determining whether the misclassification is still generated based on the assigned classification score being below a particular threshold. Every one of those limitations is recited in, or is an obvious variation of, the claims of the patent: claim 18(a)-(f) of the patent recites the identical receive-train-determine-score-request-receive-redetermine sequence, including verbatim the “assigned classification score being below a particular threshold” limitation; claim 15(d) and claim 8 of the patent recite that the synthetic dataset includes more of the particular (selected) data type than the original dataset, where that data type is determined based on the misclassification; and claim 2 of the patent recites training the model on a non-synthetic dataset and a further synthetic dataset, which is the claimed “combination of the at least one dataset and the synthetic dataset.” The instant claims differ from the patented claims only in (i) omitting the patent’s express “iterating” limitation, which merely broadens the instant claims relative to the patented claims, and (ii) reciting the subject matter in a different statutory class (method, non-transitory computer-accessible medium and system) than the corresponding patented claim, which is a matter of obvious claim-format selection given that the patent claims all three of the underlying computer-accessible medium (claims 1 and 15), method (claim 18) and “computer hardware arrangement” (claim 18(g)) embodiments. A person of ordinary skill in the art at the time of the invention would therefore have found the instant claims to be obvious variations of the patented claims, and allowing the instant claims would improperly extend the right to exclude already granted in U.S. Patent No. 11,631,032. The correspondence between the conflicting claims is set forth in the table below.
Instant Claim 41 ↔ U.S. Pat. No. 11,631,032 Claim 18
Instant Application Claims
U.S. Pat. No. 11,631,032 Claims
A method performed by a computer hardware arrangement, the method comprising:
Claim 18: A method, comprising: ... (g) using a computer hardware arrangement, iterating procedures (d)-(f) until the at least one misclassification is no longer determined during the training of the at least one model.
training at least one model on at least one dataset including a plurality of data types;
Claim 18(a)-(b): receiving at least one dataset, wherein the at least one dataset includes a plurality of data types; determining if at least one misclassification is generated during a training of at least one model on the at least one dataset;
determining at least one misclassification of one of the plurality of data types;
Claim 18(b): ... by determining if one of the data types is misclassified using at least one training model;
assigning a classification score to each of the plurality of data types after the training of the at least one model on the at least one dataset;
Claim 18(c): assigning a classification score to each of the data types after the training of the at least one model;
sending a request for at least one synthetic dataset based on the at least one misclassification;
Claim 18(d): sending a request for at least one synthetic dataset based on the misclassification;
receiving the at least one synthetic dataset;
Claim 18(e): receiving the at least one synthetic dataset;
wherein the at least one synthetic dataset includes more of a particular one of the plurality of data types than the at least one dataset, wherein the particular one of the plurality of data types is based on the at least one misclassification;
Claim 15(d): generating at least one synthetic dataset based on the misclassified at least one particular data type, wherein the at least one synthetic dataset includes more of the at least one particular data type than the at least one dataset; Claim 8: the at least one synthetic dataset includes more data samples of a selected one of the data types than the at least one dataset, wherein the selected one of the data types is determined based on the at least one misclassification.
training the at least one model on a combination of the at least one dataset and the synthetic dataset; and
Claim 18(g): using a computer hardware arrangement, iterating procedures (d)-(f) until the at least one misclassification is no longer determined during the training of the at least one model; Claim 2: train the at least one training model on at least one non-synthetic dataset and at least one further synthetic dataset.
determining if the at least one misclassification is generated during the training of the at least one model on the at least one synthetic dataset based on the assigned classification score being below a particular threshold.
Claim 18(f): determining if the at least one misclassification is generated during the training of the at least one model on the at least one synthetic dataset based on the assigned classification score being below a particular threshold;
Instant Claim 42 ↔ U.S. Pat. No. 11,631,032 Claim 18
The method of claim 41, wherein the method further comprises sending a request for the at least one synthetic dataset.
Claim 18(d): sending a request for at least one synthetic dataset based on the misclassification;
Instant Claim 43 ↔ U.S. Pat. No. 11,631,032 Claim 18
The method of claim 41, wherein the at least one synthetic dataset is based on the misclassification.
Claim 18(d): sending a request for at least one synthetic dataset based on the misclassification;
Instant Claim 44 ↔ U.S. Pat. No. 11,631,032 Claim 19
The method of claim 41, wherein the method further comprises sending a request for additional data related to a particular one of the data types.
Claim 19: The method of claim 18, wherein request includes a data request for additional data related to a particular one of the data types.
Instant Claim 45 ↔ U.S. Pat. No. 11,631,032 Claim 20
The method of claim 41, wherein the at least one dataset includes one of (i) only real data, (ii) only synthetic data, or (iii) a combination of real data and synthetic data.
Claim 20: The method of claim 19, wherein the at least one dataset includes one of (i) only real data, (ii) only synthetic data, or (iii) a combination of real data and synthetic data.
Instant Claim 46 ↔ U.S. Pat. No. 11,631,032 Claim 2
The method of claim 41, wherein the at least one model is trained on at least one non-synthetic dataset and at least one further synthetic dataset.
Claim 2: The computer-accessible medium of claim 1, wherein the computer arrangement is further configured to train the at least one training model on at least one non-synthetic dataset and at least one further synthetic dataset.
Instant Claim 47 ↔ U.S. Pat. No. 11,631,032 Claim 4
The method of claim 41, wherein the at least one dataset includes an identification of each of the plurality of data types in the at least one dataset.
Claim 4: The computer-accessible medium of claim 1, wherein the at least one dataset includes an identification of each of the data types in the at least one dataset.
Instant Claim 48 ↔ U.S. Pat. No. 11,631,032 Claim 5
The method of claim 41, wherein the method further comprises verifying an accuracy of the at least one model using at least one verification model.
Claim 5: The computer-accessible medium of claim 1, wherein the computer arrangement is further configured to verify an accuracy of the at least one training model using at least one verification model.
Instant Claim 49 ↔ U.S. Pat. No. 11,631,032 Claim 9
The method of claim 41, wherein the at least one model is a machine learning procedure.
Claim 9: The computer-accessible medium of claim 1, wherein the at least one model is a machine learning procedure.
Instant Claim 50 ↔ U.S. Pat. No. 11,631,032 Claim 1
A non-transitory computer-accessible medium having stored thereon computer-executable instructions, wherein, when a computer arrangement executes the instructions, the computer arrangement is configured to perform procedures comprising:
Claim 1: A non-transitory computer-accessible medium having stored thereon computer-executable instructions, wherein, when a computer arrangement executes the instructions, the computer arrangement is configured to perform procedures comprising:
training at least one model on at least one dataset including a plurality of data types;
Claim 1(a)-(b): receiving at least one dataset, wherein the at least one dataset includes a plurality of data types; determining if at least one misclassification is generated during a training of at least one model on the at least one dataset;
determining at least one misclassification of one of the plurality of data types;
Claim 1(b): ... by determining if one of the data types is misclassified using at least one training model;
assigning a classification score to each of the plurality of data types after the training of the at least one model on the at least one dataset;
Claim 1(c): assigning a classification score to each of the data types after the training of the at least one model;
sending a request for at least one synthetic dataset based on the at least one misclassification;
Claim 1(d): generating at least one synthetic dataset based on the at least one misclassification; Claim 18(d): sending a request for at least one synthetic dataset based on the misclassification;
receiving the at least one synthetic dataset;
Claim 18(e): receiving the at least one synthetic dataset;
wherein the at least one synthetic dataset includes more of a particular one of the plurality of data types than the at least one dataset, wherein the particular one of the plurality of data types is based on the at least one misclassification;
Claim 15(d): generating at least one synthetic dataset based on the misclassified at least one particular data type, wherein the at least one synthetic dataset includes more of the at least one particular data type than the at least one dataset; Claim 8: the at least one synthetic dataset includes more data samples of a selected one of the data types than the at least one dataset, wherein the selected one of the data types is determined based on the at least one misclassification.
training the at least one model on a combination of the at least one dataset and the synthetic dataset; and
Claim 1(f): iterating procedures (d) and (e) until the at least one misclassification is no longer determined during the training of the at least one model; Claim 2: train the at least one training model on at least one non-synthetic dataset and at least one further synthetic dataset.
determining if the at least one misclassification is generated during the training of the at least one model on the at least one synthetic dataset based on the assigned classification score being below a particular threshold.
Claim 1(e): determining if the at least one misclassification is generated during the training of the at least one model on the at least one synthetic dataset based on the assigned classification score being below a particular threshold;
Instant Claim 51 ↔ U.S. Pat. No. 11,631,032 Claim 1
The computer-accessible medium of claim 50, wherein the computer arrangement is further configured to send a request for the at least one synthetic dataset.
Claim 1(d): generating at least one synthetic dataset based on the at least one misclassification; Claim 18(d): sending a request for at least one synthetic dataset based on the misclassification;
Instant Claim 52 ↔ U.S. Pat. No. 11,631,032 Claim 1
The computer-accessible medium of claim 50, wherein the at least one synthetic dataset is based on the misclassification.
Claim 1(d): generating at least one synthetic dataset based on the at least one misclassification; Claim 18(d): sending a request for at least one synthetic dataset based on the misclassification;
Instant Claim 53 ↔ U.S. Pat. No. 11,631,032 Claim 19
The computer-accessible medium of claim 50, wherein the computer arrangement is further configured to send a request for additional data related to a particular one of the data types.
Claim 19: The method of claim 18, wherein request includes a data request for additional data related to a particular one of the data types.
Instant Claim 54 ↔ U.S. Pat. No. 11,631,032 Claim 3
The computer-accessible medium of claim 50, wherein the at least one dataset includes one of (i) only real data, (ii) only synthetic data, or (iii) a combination of real data and synthetic data.
Claim 3: The computer-accessible medium of claim 1, wherein the at least one dataset includes one of (i) only real data, (ii) only synthetic data, or (iii) a combination of real data and synthetic data. Claim 17: The computer-accessible medium of claim 15, wherein the at least one dataset includes one of (i) only real data, (ii) only synthetic data, or (iii) a combination of real data and synthetic data.
Instant Claim 55 ↔ U.S. Pat. No. 11,631,032 Claim 2
The computer-accessible medium of claim 50, wherein the at least one model is trained on at least one non-synthetic dataset and at least one further synthetic dataset.
Claim 2: The computer-accessible medium of claim 1, wherein the computer arrangement is further configured to train the at least one training model on at least one non-synthetic dataset and at least one further synthetic dataset. Claim 16: The computer-accessible medium of claim 15, wherein the computer arrangement is further configured to train the at least one training model on at least one non-synthetic dataset and at least one further synthetic dataset.
Instant Claim 56 ↔ U.S. Pat. No. 11,631,032 Claim 4
The computer-accessible medium of claim 50, wherein the at least one dataset includes an identification of each of the plurality of data types in the at least one dataset.
Claim 15(a): receiving at least one dataset including an identification of a plurality of data types in the at least one dataset; Claim 4: the at least one dataset includes an identification of each of the data types in the at least one dataset.
Instant Claim 57 ↔ U.S. Pat. No. 11,631,032 Claim 5
The computer-accessible medium of claim 50, wherein the computer arrangement is further configured to verify an accuracy of the at least one model using at least one verification model.
Claim 5: The computer-accessible medium of claim 1, wherein the computer arrangement is further configured to verify an accuracy of the at least one training model using at least one verification model.
Instant Claim 58 ↔ U.S. Pat. No. 11,631,032 Claim 9
The computer-accessible medium of claim 50, wherein the at least one model is a machine learning procedure.
Claim 9: The computer-accessible medium of claim 1, wherein the at least one model is a machine learning procedure.
Instant Claim 59 ↔ U.S. Pat. No. 11,631,032 Claim 1
A system comprising: a processor configured to:
Claim 1: ... wherein, when a computer arrangement executes the instructions, the computer arrangement is configured to perform procedures comprising: ...; Claim 18(g): using a computer hardware arrangement, iterating procedures (d)-(f) ...
train at least one model on at least one dataset including a plurality of data types;
Claim 1(a)-(b): receiving at least one dataset, wherein the at least one dataset includes a plurality of data types; determining if at least one misclassification is generated during a training of at least one model on the at least one dataset;
determine at least one misclassification of one of the plurality of data types;
Claim 1(b): ... by determining if one of the data types is misclassified using at least one training model;
assign a classification score to each of the plurality of data types after the training of the at least one model on the at least one dataset;
Claim 1(c): assigning a classification score to each of the data types after the training of the at least one model;
send a request for at least one synthetic dataset based on the at least one misclassification;
Claim 1(d): generating at least one synthetic dataset based on the at least one misclassification; Claim 18(d): sending a request for at least one synthetic dataset based on the misclassification;
receive the at least one synthetic dataset;
Claim 18(e): receiving the at least one synthetic dataset;
wherein the at least one synthetic dataset includes more of a particular one of the plurality of data types than the at least one dataset, wherein the particular one of the plurality of data types is based on the at least one misclassification;
Claim 15(d): generating at least one synthetic dataset based on the misclassified at least one particular data type, wherein the at least one synthetic dataset includes more of the at least one particular data type than the at least one dataset; Claim 8: the at least one synthetic dataset includes more data samples of a selected one of the data types than the at least one dataset, wherein the selected one of the data types is determined based on the at least one misclassification.
train the at least one model on a combination of the at least one dataset and the synthetic dataset; and
Claim 1(f): iterating procedures (d) and (e) until the at least one misclassification is no longer determined during the training of the at least one model; Claim 2: train the at least one training model on at least one non-synthetic dataset and at least one further synthetic dataset.
determine if the at least one misclassification is generated during the training of the at least one model on the at least one synthetic dataset based on the assigned classification score being below a particular threshold.
Claim 1(e): determining if the at least one misclassification is generated during the training of the at least one model on the at least one synthetic dataset based on the assigned classification score being below a particular threshold;
Instant Claim 60 ↔ U.S. Pat. No. 11,631,032 Claim 9
The system of claim 59, wherein the at least one model is a machine learning procedure.
Claim 9: The computer-accessible medium of claim 1, wherein the at least one model is a machine learning procedure.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALAN CHEN whose telephone number is (571)272-4143. The examiner can normally be reached M-F 10-7.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALAN CHEN/Primary Examiner, Art Unit 2125