Prosecution Insights
Last updated: August 17, 2026
Application No. 17/359,382

DATA PREPARATION FOR USE WITH MACHINE LEARNING

Non-Final OA §102§103
Filed
Jun 25, 2021
Priority
Nov 30, 2020 — provisional 63/119,282
Examiner
SANKS, SCHYLER S
Art Unit
2129
Tech Center
2100 — Computer Architecture & Software
Assignee
Amazon Technologies Inc.
OA Round
5 (Non-Final)
73%
Grant Probability
Favorable
5-6
OA Rounds
0m
Est. Remaining
89%
With Interview

Examiner Intelligence

Grants 73% — above average
73%
Career Allowance Rate
376 granted / 517 resolved
+17.7% vs TC avg
Strong +16% interview lift
Without
With
+15.9%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
26 currently pending
Career history
546
Total Applications
across all art units

Statute-Specific Performance

§101
2.4%
-37.6% vs TC avg
§103
46.2%
+6.2% vs TC avg
§102
16.4%
-23.6% vs TC avg
§112
34.6%
-5.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 517 resolved cases

Office Action

§102 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Reopening of Prosecution In view of the Information Disclosure Statement (IDS) filed on 06/16/2026, PROSECUTION IS HEREBY REOPENED, see MPEP 1207.04. A new ground of rejection is set forth below. To avoid abandonment of the application, appellant must exercise one of the following two options: (1) file a reply under 37 CFR 1.111 (if this Office action is non-final) or a reply under 37 CFR 1.113 (if this Office action is final); or, (2) initiate a new appeal by filing a notice of appeal under 37 CFR 41.31 followed by an appeal brief under 37 CFR 41.37. The previously paid notice of appeal fee and appeal brief fee can be applied to the new appeal. If, however, the appeal fees set forth in 37 CFR 41.20 have been increased since they were previously paid, then appellant must pay the difference between the increased fees and the amount previously paid. A Supervisory Patent Examiner (SPE) has approved of reopening prosecution by signing below: /MICHAEL J HUNTLEY/Supervisory Patent Examiner, Art Unit 2129 Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claim(s) 1-2 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Yin (CN107169575A1). Regarding claim 1, Yin teaches a computer-implemented method, comprising: obtaining from a machine-learning (ML) user interface, a textual representation of an ML training data-preparation graph, the graph comprising at least a first node identifying a data source comprising data to be prepared for training an ML model and a second node identifying a processing action to perform on the data to be prepared for training the ML model (Figure 1, see below, and Page 1 “A method for modeling a visual machine learning training model includes the following steps: S1. Select a predetermined graphical algorithm component and drag it to the design area to establish data flow between the algorithms in the graphical algorithm component to generate a process description language. [obtaining from a machine-learning user interface, a textual representation of an ML training data-preparation graph]) … Wherein, in step S1, the graphical algorithm component is formed by encapsulating a predetermined algorithm. For example, based on Canvas technology, SmartML (data modeling language SmartML is based on the JSON format, including the establishment of dataNode, query, mapping, outputTable, sql, and partition six child nodes under the root. Among them, the dataSource node is used to indicate Where the extracted data comes from. [at least a first node identifying a data source comprising data to be prepared for training an ML model]” → Page 2, “The data preprocessing component is used by the user to select a data preprocessing component for preprocessing the data in the machine learning training model; a text analysis component for the user to choose to establish a text analysis component for text analysis in a machine learning training model; A machine learning component for use by the user to establish a machine learning component for machine learning in a machine learning training model; [a second node identifying a processing action to perform on the data to be prepared for training the ML model]”) parsing the textual representation to generate computer-executable instructions (Page 1, “S2: Parse the process description language” [parsing the textual representation to generate computer-executable instructions]), the computer-executable instructions comprising a first set of computer-executable instructions generated based on a first portion of the textual representation (Page 1, “create a corresponding learning component according to the node class name and attributes, and generate a corresponding Spark learning pipeline” as applied to the first node, [a first set of computer-executable instructions generated based on a first portion of the textual representation]) and a second set of computer-executable instructions generated based on a second portion of the textual representation (Page 1, “create a corresponding learning component according to the node class name and attributes, and generate a corresponding Spark learning pipeline”, as applied to the second node [a second set of computer-executable instructions generated based on a second portion of the textual representation]) and executing the computer-executable instructions to generate an output (Page 6, “3 For the preprocessed data, drag and drop from the right to select the algorithm for analysis, and implement the model training process to obtain training results.” [to generate an output]), the output comprising at least a modified version of at least a portion of the data comprised in the data source, the modified version of the data generated in accordance with the processing action identified by the second node of the graph (Page 6, “3 For the preprocessed data, drag and drop from the right to select the algorithm for analysis, and implement the model training process to obtain training results.” [the output comprising at least a modified version of at least a portion of the data comprised in the data source, the modified version of the data generated in accordance with the processing action identified by the second node of the graph]” – The output can either be the output of the machine learning model or the output of the pre-processing step which prepares data to be input into the machine learning model. Either interpretation fits the claim limitation because both constitute a modified version of the data in the data source which is “in accordance” with the second node’s actions). PNG media_image1.png 524 902 media_image1.png Greyscale Yin, Figure 1 Regarding claim 2, Yin teaches all of the limitations of claim 1, wherein the first portion of the textual representation is associated with the first node of the graph and the second portion of the textual representation is associated with the second node of the graph (Page 1, “create a corresponding learning component according to the node class name and attributes, and generate a corresponding Spark learning pipeline” as applied to the first node, [the first portion of the textual representation is associated with the first node of the graph]” and Page 1, “create a corresponding learning component according to the node class name and attributes, and generate a corresponding Spark learning pipeline”, as applied to the second node [the second portion of the textual representation is associated with the second node of the graph]). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention. Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yin (CN107169575A1) in view of Code Ship (The Code Ship, “A guide to Python’s function decorators”, January 6, 2014). Regarding claim 3, Yin teaches all of the limitations of claim 1, but does not teach analyzing the computer-executable instructions to determine a decorator to be added to the first set of computer-executable instructions or the second set of computer- executable instructions; and adding the decorator to the first set of computer-executable instructions or the second set of computer-executable instruction before executing the computer-executable instructions to generate the output. It is known that adding decorators to functions can alter their functionality without having to alter the function itself, (The Code Ship, “In the context of design patterns, decorators dynamically alter the functionality of a function, method or class without having to directly use subclasses. This is ideal when you need to extend the functionality of functions that you don't want to modify.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include analyzing the computer-executable instructions to determine a decorator to be added to the first set of computer-executable instructions or the second set of computer-executable instructions; and adding the decorator to the first set of computer-executable instructions or the second set of computer-executable instruction before executing the computer-executable instructions to generate the output in Yin in order to allow the functions in Xu to be appropriately modified for a specific task without having to alter the core framework of the function. Claim(s) 4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yin (CN107169575A1) in view of Smith (US20070040094A1). Regarding claim 4, Yin teaches all of the limitations of claim 1, but does not teach obtaining information indicating a user selected mode usable to determine a quantity of modified version of the data to include in the output. Smith teaches obtaining information indicating a user selected mode usable to determine a quantity of modified version of the data to include in the output (¶115, “Create new categorical variables from continuous variables by splitting the numeric values into a number of bins. This is useful, for example, if you have an age column; rather than including age as a continuous variable in your models, it might be more beneficial to split it into ranges that make sense demographically”, Figure 24, “Split” determines some quantity of modified data to include in the output bins). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include obtaining information indicating a user selected mode usable to determine a quantity of modified version of the data to include in the output in the program of Yin in order to allow the data to be altered/split into ranges that make more sense for the particular application. Claim(s) 5-7, 10-15, and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yin (CN107169575A) in view of Singaraju (US20200344185A1). Regarding claim 5, Yin teaches a system performing the method executed by the claimed processor (see rejection of claim 1). Yin does not explicitly disclose wherein the system comprises one or more processors and memory that stores computer-executable instructions that are executable by the one or more processors to cause the system to perform the claimed methodology. Singaraju teaches where the system comprises one or more processors and memory that stores computer-executable instructions that are executable by the one or more processors to cause the system to perform workflow operations (¶12, “In various embodiments, a system is provided that includes one or more data processors and a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein.” [one or more processors; and memory that stores computer-executable instructions that are executable by the one or more processors to cause the system to]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize a one or more processors and memory in Yin to perform operations in Yin in order to efficiently perform the operations of Yin on a computer. Regarding claim 6, Yin as modified teaches all of the limitations of claim 5. Yin does not teach wherein the textual representation of the graph comprises text, the text comprising at least a node identifier (ID) for the second node and a function name. Singaraju teaches wherein the textual representation of the graph comprises text, the text comprising at least a node identifier (ID) for the second node and a function name (Figure 3, see below, 305 shows “Name: Trim” [comprises text, the text comprising at least a node identifier (ID) for the second node] as well as “OP: Trim String” [and a function name]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to utilize text in the textual representation of the graph, the text comprising at least a node identifier (ID) for the second node and a function name in order to efficiently convert the graph structure of Yin into a form executable by a computer (¶100 of Singaraju, “Essentially, all tasks or activities to be implemented in a model are laid out in a clear structure or pipeline with discrete processes occurring at set points and clear relationships made to other tasks. If multiple tasks exist, each has at least one defined upstream (previous) or downstream (subsequent) tasks, although each task could easily have both. No task can create data that goes on to reference itself (this avoids any instance of an infinite loop). As shown in FIG. 3, a DAG based pipeline 300 may be defined as a JSON object containing an acyclic array of nodes 305, and connection 310 between the nodes. The pipeline 300 is a graph which holds the track of operations (OP) applied on a dataset (DF) (e.g., a data table) such as Resilient Distributed Datasets (RDD). The track of operations are directly connected from one node to another. This creates a sequence i.e. each node is in linkage from earlier to later in the appropriate sequence, and each node automatically identifies outputs from a previous node as inputs and output of a current node as input for a next node. Each node within the sequence adds a new field to the DF, which is moved from one node to another as the workflow progresses through the pipeline 300. Moreover, the pipeline is defines such that there is no cycle or loop available. Once a transformation takes place it cannot return to its earlier position. Due to the acyclic nature of the nodes and connectivity thereof, the cluster-computing framework is capable of automatically identifying cyclic dependencies and rejects the pipeline during a pipeline detection stage. FIG. 4 shows an invalid pipeline 400 with a cyclic dependency.”) PNG media_image2.png 578 796 media_image2.png Greyscale Singaraju, Figure 3 Regarding claim 7, Yin as modified teaches all of the limitations of claim 6, wherein parsing the computer-executable instructions comprises: determining the function name comprised in the text; and locating an executable function based on the function name comprised in the text (see ¶104 snippet from Singaraju below. “trim” and ¶101, “Each node 305 has a stage definition comprising: a “Name”, a class to be invoked via an introspective function such as Java Reflection, parameters needed by the class, and one or more operations (OP) performed on a given data set.” [determining the function name comprised in the text and locating an executable function based on the function name comprised in the text]). PNG media_image3.png 410 610 media_image3.png Greyscale Singaraju, ¶104 Regarding claim 10, Yin as modified teaches all of the limitations of claim 5, wherein the processing to be performed on the data corresponds to at least one predefined transform function that, when executed, is to modify the data, the at least one predefined transform function comprising the computer-executable instructions stored and executed to generate the ML training data from the data (Yin, “S2, analyzing the process description language, establishing corresponding learning component according to the node type name and attribute [at least one predefined transform function that, when executed, is to modify the data], and generating corresponding learning Spark pipeline; S3, the study pipeline submitted for model training on the Spark cluster [the at least one predefined transform function comprising the computer-executable instructions stored and executed to generate the ML training data from the data]”) Regarding claim 11, Yin as modified teaches all of the limitations of claim 5, wherein the graph further comprises another node representing a data source storing the data, the obtained textual representation of the graph identifying the data source storing the data (Singaraju, Figure 3, “Name: Extract Data”, “OP: Extract from DB” [another node representing a data source storing the data, the obtained textual representation of the graph identifying the data source storing the data]). Regarding claim 12, Yin as modified teaches all of the limitations of claim 11, wherein the computer-executable instructions that are executable by the one or more processors to further cause the system to: retrieve the data from the data source based on determining a location of the data source from a portion of the textual representation that identifies the data source storing the data, and wherein the computer-executable instructions use the retrieved data to generate the ML training data (Singaraju, Figure 3, “OP: Extract from DB”). Regarding claims 13-15, Yin as modified according to claims 5-7 teaches the non-transitory computer-readable storage medium of claims 13-15. Regarding claim 17, Yin as modified teaches all of the limitations of claim 13, wherein the obtained textual representation of the ML training data-preparation graph identifies a portion of the ML graph corresponding to the one or more transforms usable to prepare the data for ML training (Yin, “S2, analyzing the process description language, establishing corresponding learning component according to the node type name and attribute [identifies a portion of the ML graph corresponding to the one or more transforms usable to prepare the data], and generating corresponding learning Spark pipeline; S3, the study pipeline submitted for model training on the Spark cluster [for ML training]”) Claim(s) 9, 16, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yin (CN107169575A1) in view of Singaraju (US20200344185A1), further in view of Sainani (US20190034767A1). Regarding claim 9, Yin as modified teaches all of the limitations of claim 5. Yin as modified does not teach wherein the computer-executable instructions that are executable by the one or more processors to further cause the system to transmit the ML training data to a client computing device, at least a portion of the ML training data to be displayed in the ML UI. Sainani teaches transmitting the ML training data to a client computing device, at least a portion of the ML training data to be displayed in the ML UI (Figure 6B, see below, where data is displayed [transmitting the ML training data to a client computing device] and is displayed in a UI [at least a portion of the ML training data to be displayed in the ML UI]). PNG media_image4.png 430 766 media_image4.png Greyscale Sainini, Figure 6B It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Yin to include causing the system to transmit the ML training data to a client computing device, at least a portion of the ML training data to be displayed in the ML UI in order to allow the user to view training data. Regarding claim 16, Yin as modified according to claim 9 teaches the computer readable storage medium of claim 16. Regarding claim 20, Yin as modified according to claim 16 teaches the limitations of claim 20 because the transmission of ML training data can be considered the transmission of a message comprising the ML training data. Claim(s) 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Yin (CN107169575A1) in view of Singaraju (US20200344185A1) as applied to claim 13, further in view of Code Ship (The Code Ship, “A guide to Python’s function decorators”, January 6, 2014). Regarding claim 19, Yin as modified teaches all of the limitations of claim 13, but does not teach analyzing the computer-executable instructions to determine a decorator to be added to the first set of computer-executable instructions or the second set of computer- executable instructions; and adding the decorator to the first set of computer-executable instructions or the second set of computer-executable instruction before executing the computer-executable instructions to generate the output. It is known that adding decorators to functions can alter their functionality without having to alter the function itself, (The Code Ship, “In the context of design patterns, decorators dynamically alter the functionality of a function, method or class without having to directly use subclasses. This is ideal when you need to extend the functionality of functions that you don't want to modify.”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to include analyzing the computer-executable instructions to determine a decorator to be added to the first set of computer-executable instructions or the second set of computer-executable instructions; and adding the decorator to the first set of computer-executable instructions or the second set of computer-executable instruction before executing the computer-executable instructions to generate the output in Yin in order to allow the functions in Xu to be appropriately modified for a specific task without having to alter the core framework of the function. Allowable Subject Matter Claims 8 and 18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Regarding claims 8 and 18, the claims require the receiving of a message from a client computing device in obtaining the claimed textual representation of the one or more transforms. Per the claims, the message includes a node identification, an identification of the transforms, and a portion of the data to transform. Yin does not describe this process. Furthermore, the prior art does not establish a prima facie case of obviousness for receiving such a message in Yin. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Wang (US20180373961A1) discloses a method for reading a graphical topology and deploying functions according to the graphical topology. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SCHYLER S SANKS whose telephone number is (571)272-6125. The examiner can normally be reached 06:30 - 15:30 Central Time, M-F. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael Huntley can be reached at (303) 297-4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SCHYLER S SANKS/Primary Examiner, Art Unit 2129
Read full office action

Prosecution Timeline

Show 7 earlier events
Jul 11, 2025
Non-Final Rejection mailed — §102, §103
Oct 14, 2025
Response Filed
Dec 04, 2025
Final Rejection mailed — §102, §103
Jan 29, 2026
Response after Non-Final Action
Mar 04, 2026
Notice of Allowance
May 04, 2026
Response after Non-Final Action
May 18, 2026
Response after Non-Final Action
Jul 27, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12693058
Expansion Valve Position Detection in Refrigeration System
3y 4m to grant Granted Jul 28, 2026
Patent 12682275
LEARNING MODEL APPLYING SYSTEM, A LEARNING MODEL APPLYING METHOD, AND A PROGRAM
5y 0m to grant Granted Jul 14, 2026
Patent 12681743
Virtual Machine Managing System Using Snapshot
3y 11m to grant Granted Jul 14, 2026
Patent 12675670
Offline Primitive Discovery For Accelerating Data-Driven Reinforcement Learning
3y 4m to grant Granted Jul 07, 2026
Patent 12670404
METHOD AND SYSTEM FOR TRAINING A NEURAL NETWORK MODEL USING KNOWLEDGE DISTILLATION
4y 9m to grant Granted Jun 30, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
73%
Grant Probability
89%
With Interview (+15.9%)
2y 10m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 517 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month