Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAIL ACTION
Priority
Acknowledgment is made of applicant's claim for foreign priority under 35 U.S.C. 119(a)-(d). The certified copy has been placed of record in the file.
Information Disclosure Statement
The information disclosure statement (IDS) was submitted on 1/25/2025 and 9/1/2025. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-10 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claimed invention is directed to non-statutory subject matter because the claim(s) as a whole, considering all claim elements both individually and in combination, do not amount to significantly more than an abstract idea. As summarized in the 2019 Revised Patent Subject Matter Eligibility Guidance, examiners must perform a Two-Part Analysis for Judicial Exceptions.
Step 1
In Step 1, it must be determined whether the claimed invention is directed to a process, machine, manufacture or composition of matter. The instant invention encompasses two sets of claims: a method in claims 1-8 (i.e., a process) and a system in claims 9-10 (i.e., a manufacture). All claims are directed to one of the four statutory categories and meet the requirements of step 1.
Step 2A
Prong One
The claimed invention is directed to an abstract idea without significant more. The instant invention is broadly directed to “generating training data for fine-tuning of a Machine learning model”. Claim 1 recites the following (with emphasis added):
Claim 1: A method of generating training data for fine-tuning of a Machine Learning (ML) model, the method comprising:
generating one or more natural language interpretations of a dataset corresponding to one or more parameters associated with configuration of the dataset, wherein each of the one or more natural language interpretations is text-based;
collating the one or more natural language interpretations of the dataset corresponding to one or more parameters, to generate a combined natural language interpretation of the dataset, wherein the combined natural language interpretation is text-based;
generating a conceptual explanation of the dataset, based on the combined natural language interpretation of the dataset; and
assigning one or more labels to each sub-dataset of the dataset, based on the conceptual explanation of the dataset, to generate training data for fine-tuning of the ML model, wherein the dataset comprises a plurality of sub-datasets.
The bold portions of claim 1 encompass the abstract idea, which is also encompassed by the dependent claims 2-8, and substantially also encompassed by claims 9-10.
Claims 1 and 9 recite the steps to generate training data for fine-tuning of a Machine learning model including a natural language processing. These limitations, when given their broadest reasonable interpretation, are directed to certain performing of organizing human activity and mental processes, which is abstract idea.
Prong Two
This judicial exception is not integrated into a practical application because mere instruction to implement on computers (i.e. memory or processors in claim 9) or a computer model (machine learning model here in claim 1), or merely using computers as a tool to perform the abstract idea, adding insignificant extra solution activity, and/or generally linking the use of the abstract idea to a technological environment for field of use is not considered integration into a practical application. Claim 1 recites using natural language process to generate training data for fine-tuning of a Machine learning model. Using training data to a trained machine-learning model or natural language model is a generic feature of natural language process, which does not represent a technological improvement. The using of the computer and natural language process does not add improvement to the functioning of a computer or to any other technology field, which failed to enable the abstract idea to integrate into a practical application. The claims are drafted in a result-oriented fashion, without the requisite specificity needed to provide a nonabstract technological solution. The computing system and natural language process are directed to the components of a system amount to merely field of use type limitations and/or extra solution activity to implement the abstract idea as presented.
Step 2B
Step 2B in the analysis requires us to determine whether the claims do significantly more than simply describe that abstract method. Mayo, 132 S. Ct. at 1297. We must examine the limitations of the claims to determine whether the claims contain an "inventive concept" to "transform" the claimed abstract idea into patent-eligible subject matter. Alice, 134 S. Ct. at 2357 (quoting Mayo, 132 S. Ct. at 1294, 1298). The transformation of an abstract idea into patent-eligible subject matter "requires 'more than simply stat[ing] the [abstract idea] while adding the words 'apply it."' Id. (quoting Mayo, 132 S. Ct. at 1294) (alterations in original). "A claim that recites an abstract idea must include 'additional features' to ensure 'that the [claim] is more than a drafting effort designed to monopolize the [abstract idea].'" Id. (quoting Mayo, 132 S. Ct. at 1297) (alterations in original). Those "additional features" must be more than "well-understood, routine, conventional activity." Mayo, 132 S. Ct. at 1298.
The present claims include the additional elements other than the abstract idea which include a processor, memory, and machine learning model (in claim 1 and 9). These additional elements are merely conventional computer and computer model. Any potentially technical aspects of the claims are well-known generic computer components performing conventional functions (e.g., a processor performing a mental process). The present claims have been analyzed both individually and in combination and, the instant claims do not provide any improvement of the functioning of the computer or improvement to computer technology or any other technical field. There do not appear to be any meaningful limitations other than those that are well-understood, routine and conventional in the field. Thus, the present claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. Thus, the claims 1-8 are not patent eligible.
Claims 9-10 recite similar limitations of claims 1-8, thus are abstract idea and not patent eligible.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1 and 9 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vijaykeerthy et al (US 20210012156 A1) in view of Qian et al (US 20190311229 A1).
Regarding claim 1, Vijaykeerthy discloses a method of generating training data for fine-tuning of a Machine Learning (ML) model [e.g. FIG. 1 and 3; Re-train a machine learning model based on a Subset of training examples], the method comprising:
generating one or more natural language interpretations of a dataset [e.g. 102 and 106; explanations for the identified training examples] corresponding to one or more parameters [e.g. [0027-0028]; parameter configuration and used to refine the model] associated with configuration of the dataset e.g. explanations for the identified training examples], wherein each of the one or more natural language interpretations is text-based [e.g. [0026]; the input dataset may be in the form of text];
collating the one or more natural language interpretations of the dataset corresponding to one or more parameters [e.g. FIG. 2-3; model parameters; a measure of influence in selecting the subset of training examples] to generate a combined natural language interpretation of the dataset [e.g. user explanation for each of a subset of training examples], wherein the combined natural language interpretation is text-based [e.g. [0026]; the input dataset may be in the form of text; the annotations may be provided using any suitable type of user input, or combination thereof, such as audio, text, drawings, etc.];
generating a conceptual explanation of the dataset [e.g. FIG. 2-3; identified parts of the training examples in the subset], based on the combined natural language interpretation of the dataset [e.g. training a machine learning model based at least in part on the user explanations]; and
assigning one or more sub-dataset of the dataset, based on the conceptual explanation of the dataset, to generate training data [e.g. e.g. FIG. 3; Re-Train a machine learning model based at least in part on the user explanations] for fine-tuning of the ML model [e.g. a machine learning model], wherein the dataset comprises a plurality of sub-datasets [e.g. FIG. 1-3].
It is noted that Vijaykeerthy differs to the present invention in that Vijaykeerthy fails to explicitly labeling each sub-dataset of the dataset.
However, Qian teaches the well-known concept of labels to each sub-dataset of the dataset [e.g. FIG. 1-2; a set of training data] based on the conceptual explanation of the dataset [e.g. FIG. 3-4; data items labeled by user as either to be accepted or rejected].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify training a machine learning model system disclosed by Vijaykeerthy to exploit the well-known labeling the training data set technique taught by Qian as above, in order to provide implementing active learning to improve precision for learning models [See Qian; [0039]].
Regarding claim 9, this is an apparatus that includes same limitation as in claim 1 above, the rejection of which are incorporated herein.
Claim(s) 2-8 and 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Vijaykeerthy et al (US 20210012156 A1) in view of Qian et al (US 20190311229 A1) and HÖLZL et al (US 20230394331 A9).
Regarding claim 2, Vijaykeerthy and Qian further disclose identifying a domain information [e.g. Vijaykeerthy: type of service] , a context information [e.g. Qian: configuration settings], and metadata associated with the dataset [e.g. Vijaykeerthy: explanation for training examples], based on the conceptual explanation of the dataset; assigning one or more labels to the dataset [e.g. FIG. 3-4; data items labeled by user as either to be accepted or rejected]. but Vijaykeerthy and Qian fail to disclose detail of assigning label to the dataset,
However, HÖLZL teaches the well-known concept of assigning one or more labels to the dataset [e.g. FIG. 1-2; annotate data], based on the domain information [e.g. treatments of patients], the context information [e.g. data being semantically annotated], and the metadata [e.g. FIG. 1; The data sets DS are semantically annotated based on a semantic representation SR ] associated with the dataset.
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify training a machine learning model system disclosed by Vijaykeerthy to exploit the well-known labeling the training data set technique taught by Qian and the well-known concept of annotating data set technique taught by HÖLZL as above, in order to provide implementing active learning to improve precision for learning models [See Qian; [0039]] and the highest prediction quality derived from the trained machine learning [See HÖLZL; abstract and [0016]].
Regarding claim 3, Vijaykeerthy, Qian and HÖLZL further disclose the one or more parameters associated with the configuration of the dataset comprise: a structure [e.g. HÖLZL: a tree structure] associated with the dataset, one or more dependencies associated with the dataset [e.g. HÖLZL: FIG. 2; [0037]] , an architecture associated with the dataset [e.g. HÖLZL: FIG. 2], and a formatting associated with the dataset [e.g. Vijaykeerthy: FIG. 1 and 3; Qian: FIG. 1; HÖLZL: FIG. 2].
It would have been obvious to a person of ordinary skill in the art before the effective filing date of the invention to modify training a machine learning model system disclosed by Vijaykeerthy to exploit the well-known labeling the training data set technique taught by Qian and the well-known concept of annotating data set technique taught by HÖLZL as above, in order to provide implementing active learning to improve precision for learning models [See Qian; [0039]] and the highest prediction quality derived from the trained machine learning [See HÖLZL; abstract and [0016]].
Regarding claim 4, Vijaykeerthy, Qian and HÖLZL further disclose at least one of: functions, methods, and classes associated with the dataset [e.g. Vijaykeerthy: FIG. 1 and 3; [0037-0038]; class label; Qian: FIG. 1; HÖLZL: FIG. 2]; one or more files within the dataset; one or more directories within the dataset; and a core task associated with the dataset [HÖLZL: FIG. 2; medical data; treatment data].
Regarding claim 5, Vijaykeerthy, Qian and HÖLZL further disclose one or more variable dependencies, functional dependencies, and module dependencies [e.g. HÖLZL: FIG. 2; [0037]].
Regarding claim 6, Vijaykeerthy, Qian and HÖLZL further disclose an algorithmic or design pattern, data structures, exceptions, and Input and Output values [e.g. Vijaykeerthy: FIG. 1 and 3; Qian: FIG. 1; [0016]; efficiently learn ER algorithms; HÖLZL: FIG. 2].
Regarding claim 7, Vijaykeerthy, Qian and HÖLZL further disclose the formatting associated with the dataset comprises at least one of: a style, an indentation, and comments associated with the dataset [e.g. Vijaykeerthy: FIG. 1 and 3; Qian: FIG. 1; efficiently learn ER algorithms; HÖLZL: FIG. 2].
Regarding claim 8, Vijaykeerthy, Qian and HÖLZL further disclose the conceptual explanation of the dataset comprises at least one of: a high-level design associated with the dataset; and a low-level design associated with the dataset [e.g. HÖLZL: FIG. 2; hierarchical levels of the respective trees].
Regarding claim 10, this is an apparatus that includes same limitation as in claim 2 above, the rejection of which are incorporated herein.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
BAHRAMI et al (US 20230106226 A1).
BURSZTYN et al (US 20250124235 A1).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ZHUBING REN whose telephone number is (571)272-2788. The examiner can normally be reached Monday-Friday 9am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at 571-272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ZHUBING REN/ Primary Examiner, Art Unit 2658