Prosecution Insights
Last updated: October 04, 2026
Application No. 19/094,484

Kaplan-Meier Digitizer for Automating Kaplan-Meier Curve Analysis

Non-Final OA §103
Filed
Mar 28, 2025
Priority
Apr 01, 2024 — provisional 63/572,645
Examiner
LIU, GORDON G
Art Unit
Tech Center
Assignee
Bristol-Myers Squibb Company
OA Round
1 (Non-Final)
83%
Grant Probability
Favorable
1-2
OA Rounds
7m
Est. Remaining
98%
With Interview

Examiner Intelligence

Grants 83% — above average
83%
Career Allowance Rate
581 granted / 701 resolved
+22.9% vs TC avg
Moderate +15% lift
Without
With
+14.8%
Interview Lift
resolved cases with interview
Fast prosecutor
2y 2m
Avg Prosecution
41 currently pending
Career history
720
Total Applications
across all art units

Statute-Specific Performance

§101
7.1%
-32.9% vs TC avg
§103
77.3%
+37.3% vs TC avg
§102
3.4%
-36.6% vs TC avg
§112
2.7%
-37.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 701 resolved cases

Office Action

§103
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Claims 1-26 are pending under this Office action. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-7, 11-20 and 24-26 are rejected under 35 U.S.C. 103 as being unpatentable over Lahmann. etc. (US 20170351708 A1) in view of Witte, etc. (US 20180253873 A1), further in view of Bekas, etc. (US 20190130614 A1) and Xiao, etc. (US 20230177682 A1) and Xiao. etc. (US 20230177682 A1). Regarding claim 1, Lahmann teaches that a computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations (See Lahmann: Fig. 2, and [0168], “FIG. 2 depicts a computer system 200 of an embodiment of the invention that is configured for extracting data from a scatter plot 218. The extraction of data from a scatter plot may be advantageous as scatter plots are commonly used. The extraction of data from scatter plots has often been reported to be difficult due to problems to correctly identify the data points and series to which a particular data point belongs. Embodiments of the invention may allow the extraction of data from scatter plots in an efficient, error robust and accurate manner”) comprising: receiving an input image containing a graphical Kaplan-Meier (KM) plot of one or more KM curves, the graphical KM plot having an x-axis and a y-axis orthogonal to the x-axis (See Lahmann: Fig. 1, and [0160], “In a first step 102, the image analysis logic receives a digital image of a scatter plot. For example, the image analysis logic can read a JPEG RGB image that depicts a scatter plot from a storage medium or from a webpage that comprises the JPEG image. Alternatively, the image analysis logic can be coupled to an image acquisition system, e.g. a camera or a scanner, and receive the image from the image acquisition system. Then, the received image can optionally be processed for transforming the RGB image into a binary digital image. Alternatively, the digital image of the scatter plot can already be received in the form of a binary scatter plot imager”); processing the input image to convert the graphical KM plot into a three-dimensional (3D) array representing a color of each pixel from the graphical KM plot by a respective 3D vector (See Lahmann: Figs. 4A-D, and [0181], “In a step depicted in FIG. 4A, adjacent pixel groups of pixels which are similar to each other are identified, e.g. by means of a connected-component analysis, as “pixel sets” 402. For example, the pixels of the axes or of one axes may correspond to one pixel set. The pixels of each character of an axis label or of the title may correspond to one pixel set. The pixels of each symbol representing an isolated data point may correspond to a respective single pixel set. The pixels of each cluster of overlapping symbols of multiple data points may correspond to a respective single pixel set. Thus, the pixel sets or “blobs” depicted in FIG. 4A may represent data points, sets of overlapping data points, artifacts, labels, axes, characters and other object types that might cause errors”; [0203], “An inner morphological gradient filter, performed by taking a morphologically dilated image minus the original image, is applied 34 to each red-green-blue (RGB) color channel of the original image to produce three new single-channel (grayscale) images, a, b, and c. A composite grayscale image, x, is then computed from a, b, and c by selecting the maximum pixel value at each pixel coordinate in an image from images a, b, and c and storing the selected maximum pixel value into x. One or more optimal threshold computations (such as the global statistical mean and the standard deviation of pixel intensities within x), are performed 36 on x to produce a binary image featuring the contours of the graphical objects in the original image. The collection of individual connected components in the image is computed and each of these elements is used to segment, locate, and label the set of data components in the image as outlined below”. Note that the color RGB is mapped to the 3D vector per pixel, and the digital color image pixels are mapped to arrays); processing the 3D array to generate (See Lahmann: Fig. 4A-D, and [0024], “In case one or more derivative images, e.g. a binary image, is created, the template matching can be performed on a derivative image (which is typically faster) or on the original image (which is typically more accurate as it may comprise more information than a derivative image). Depending on the form of the received image and other factors (e.g. whether high accuracy or high performance is preferred), different kinds of pixel set detection algorithms, template matching approaches and/or data series assignment approaches can be applied”): a black pixel matrix mask by converting all white pixels to pure white and all other pixels to pure black (See Lahmann: Figs. 4A-D, and [0030], “The identification of pixel sets based on similar adjacent pixels can be performed e.g. by means of a connected component analysis which may be performed e.g. on an original multi-channel scatter plot image. For instance, the connected component analysis may be performed on an RGB scatter plot image. Alternatively, the identification of pixel sets based on similar adjacent pixels can be performed on a derivative image, e.g. a grayscale or binary black-white image”; and [0048], “According to embodiments, binary (black-and-white) images are generated as an intermediate step to the extraction of the connected components, i.e. pixel sets, for reducing the color information. Then, all white, repetitive black (depending on the definition) “pixel blobs” are used as “masks” which identify respective pixel sets in the multi-channel image which are further processed for extracting the templates and for performing the template mapping”); and a colored pixel matrix mask by converting all black, grey, and white pixels to pure white ; (See Lahmann: Figs. 4A-D, [0048], “According to embodiments, binary (black-and-white) images are generated as an intermediate step to the extraction of the connected components, i.e. pixel sets, for reducing the color information. Then, all white, repetitive black (depending on the definition) “pixel blobs” are used as “masks” which identify respective pixel sets in the multi-channel image which are further processed for extracting the templates and for performing the template mapping””) processing the black pixel matrix mask to identify pixel coordinates that define the x-axis and the y-axis of the graphical KM plot (See Lahmann: Fig. 7, and [0203], “An inner morphological gradient filter, performed by taking a morphologically dilated image minus the original image, is applied 34 to each red-green-blue (RGB) color channel of the original image to produce three new single-channel (grayscale) images, a, b, and c. A composite grayscale image, x, is then computed from a, b, and c by selecting the maximum pixel value at each pixel coordinate in an image from images a, b, and c and storing the selected maximum pixel value into x. One or more optimal threshold computations (such as the global statistical mean and the standard deviation of pixel intensities within x), are performed 36 on x to produce a binary image featuring the contours of the graphical objects in the original image. The collection of individual connected components in the image is computed and each of these elements is used to segment, locate, and label the set of data components in the image as outlined below”; and [0207], “Once all data points in the plot image are identified, a bounding box is initialized 44 enclosing all elements of the data set. In some embodiments, the bounding box is deformed so that its edges reside on the vertical and horizontal axes lines of the plot image. The axes are then identified 46 as the line segments representing the edges of the bounding box. The textual labels, including series labels, chart title(s), axes range values, etc. are identified and extracted using optical character recognition (OCR)”); cropping the colored pixel matrix mask based on the identified pixel coordinates that define the x-axis and the y-axis of the graphical KM plot (See Lahmann: Fig. 7, and [0207], “Once all data points in the plot image are identified, a bounding box is initialized 44 enclosing all elements of the data set. In some embodiments, the bounding box is deformed so that its edges reside on the vertical and horizontal axes lines of the plot image. The axes are then identified 46 as the line segments representing the edges of the bounding box. The textual labels, including series labels, chart title(s), axes range values, etc. are identified and extracted using optical character recognition (OCR)”); processing the cropped colored pixel matrix mask to segment the colored pixels from the cropped colored pixel matrix mask into respective groups of clustered pixels, each respective group of clustered pixels associated with a different respective color (See Lahmann: Fig. 4A-D, and [0137], comparing each of the templates with pixels of a target image, the target image being the received scatter plot image or the derivative of the received scatter plot image, for identifying positions of matching templates, a matching template being a template whose degree of similarity to pixels of the target image exceeds a similarity threshold”; and [0162], “In step 106, the image analysis logic analyzes the identified pixel sets in order to generate a plurality of templates. The template generation may involve the generation of template candidates from which the finally used templates are selected in one or more filtering steps as described, for example, with reference to FIG. 4. Each template is a pixel structure that depicts exactly one data point symbol, e.g. a red triangle or a black circle. Thus, each template and respective data point symbol represents a respective data series, whereby all data points of a particular data series are assumed to be represented in the plot with the respective symbol of that series”. Note that the plurality of templates is mapped to the respective groups of clustered pixels, each respective group of clustered pixels associated with a different respective color); processing each respective group of clustered pixels that represents a corresponding KM curve of the one or more KM curves of the graphical KM plot to generate a respective digitized representation of the corresponding KM curve (See Lahmann: Fig. 4A-D and 7, and [0206], “The identified set of data points are segmented 42 into different data series each including a plurality of data points, based, for instance, on the locations, spacings, coloring, patterns, and/or shapes of the image elements they represent. For instance, if a line chart plot image contained three lines of different colors, red, blue, and yellow, the digitization system segments the data into three separate series, with data sets corresponding to each line based on color. Similarly, if a scatter plot image contained two types of data point elements, circles and diamonds, the digitization system segments the data into two separate series, with one data set corresponding to all circle elements of the plot image and one data set corresponding to all diamond elements of the plot image. As discussed below, the different series are identified with distinct markers and are separated into partially or wholly distinct data sets in the data grid”, Note that the segment data point in series is the digitized representation of that curve/series, and it is mapped to the KM curves); and generating a digitized KM plot based on the identified pixel coordinates (See Lahmann: Fig. 7, and[0014], “returning the identified data points and the data series to which it is assigned”; [0208], “All extracted components of the plot image, including data points, axes markers, and textual labels (series labels, chart title(s), axes range values, etc.) are visually presented in the user interface 50, as specified below in steps 52-56”; and [0209], “The identified data points are marked 52, and the identified axes are marked 54 with polygons, crosshairs, or lines on the canvas overlaying the plot image. The textual labels, including series labels, chart title(s), axes range values, etc., are visually presented in the user interface 56 such that the user may manipulate these elements”. Note that the returned identified data points are mapped to the digitized KM plot) that define the x-axis and the y-axis of the graphical KM plot and the respective digitized representation generated for each corresponding KM curve of the one or more KM curves of the graphical KM plot. However, Lahmann fails to explicitly disclose that processing the black pixel matrix mask to identify pixel coordinates that define the x-axis and the y-axis of the graphical KM plot; and that define the x-axis and the y-axis of the graphical KM plot and the respective digitized representation generated for each corresponding KM curve of the one or more KM curves of the graphical KM plot. However, Witte teaches that processing the black pixel matrix mask to identify pixel coordinates that define the x-axis and the y-axis of the graphical KM plot (See Witte Figs. 4-5, and [0034], “The processing of the image may include a number of subprocesses. In particular, the stage may include: 1) detecting and removing horizontal and vertical grid lines from the image in a known manner; 2) optionally removing isolated noise by detecting and removing connected components below some number of connected pixels; 3) optionally smoothing the image in the horizontal direction using, for example, an exponential or Gaussian smoothing function; 4) optionally extracting measures of local image orientation and structure using, for example, the eigenvalues and eigenvectors of the local gradient structure tensor at each point in the image; and 5) optionally enhancing the curves in the image by, for example, simple color separation if each curve is a different color or, for example, by performing a local Hough transform or anisotropic diffusion at each point in a binary or gray-scale image”; Figs. 2-3and [0030], “Prior to describing the digital well log vectorization process, a raster well log image and the nomenclature for the regions in the raster well log image are described with reference to FIG. 1. As shown, the raster well log image may have a Header region, a Tool Configuration region, a Miscellaneous region and a Trailer region that each contain meta-information about the well, well bore, measurement devices and location. The digital well log vectorization process described below may be used for processing of the Upper Sections (US), Lower Section (LS) and the Log Sections. The Upper and Lower Sections contain scale information for the well log measurements. The Log Section contains the raster image of the well log curves. In the other figures described below, the Log sections only are shown. For example, FIGS. 2 and 3 illustrate two examples of a raster well log image containing two tracks and multiple well log curves. A left hand side of FIG. 2 is an example of a raster well log image containing two tracks and multiple well log curves. Note that in FIG. 2, the ‘horizontal’ and ‘vertical’ grid lines are neither horizontal nor vertical, indicating rotation and possible additional horizontal skew or stretch of the image. Also note the variable image quality, including a dark splotch in the lower part of the image. The right hand side of FIG. 2 shows a portion of the image in the left hand side of FIG. 2 after well-known warping the image to rectify the grid lines. FIG. 3 shows another raster well log image with two tracks, each containing multiple curves. Note the vertical scale between the tracks. This scale is used to convert from vertical pixels to measured depth along the well bore”; and fig. 6, and [0027]. “A further advantage of the system and method may be that the image-guided shortest path algorithm can also be used to improve the pre-processing step of image rectification, by automatically tracing the grid lines and track edges of the raster log image. These extracted grid lines should be strictly horizontal or vertical and intersect at right angles. Detected deviations from these conditions form the basis for accurately warping the image to remove distortions due to poor scanning or damage to the original paper records. Improved image rectification results in more accurate values for the extracted digital well log curves”; [0032], “Each of the stages of the method shown in FIGS. 4A and 4B may be implemented by the system implementations shown in FIGS. 16A and 16B or by other means known to those skilled in the art. Furthermore, the processes shown in FIGS. 4A and 4B can be implemented on specialized hardware and hardware/software designed to perform the well log vectorization that is substantially more than a general computer system since the specialized hardware and hardware/software must be specially programmed or designed to be even able to perform the processes of the well log vectorization. Returning to FIGS. 4A and 4B, the method 400 may include one or more stages that are each in a processing pipeline. Although the stages are shown in a particular order in FIGS. 4A and 4B, the order of the stages may be reordered without departing from the scope of the disclosure. In a first stage, the method may load and display a raster log image (402), an example of which is shown in FIGS. 2-3. Once the raster log image is loaded into the well log vectorizer component (or the computer system), the method may perform a process of identifying and registering the raster log image components (404). For example, during this process of identifying and registering the raster log image components, the method may vertically and horizontally register the raster log image components as shown in FIGS. 5 and 6. FIG. 5 is an example of a raster well log image track that has been vertically and horizontally registered. In FIG. 5, a set of green horizontal lines (horizontal lines with periodic triangles along the line) correspond to specific measured depths (MD) of the well log such as MD=0 and MD=2000 in the example in FIG. 5. The green vertical lines (lines with squares periodically along the line) indicate the vertical edges of the track and correspond to specific measured physical values for each curve. FIG. 6 shows a raster well log image track with multiple vertical grid lines extracted (red vertical lines). Although in this image the vertical grid lines are indeed close to vertical, that is often not the case and the image can be warped to force the lines to be vertical, thus improving the image rectification prior to well log curve extraction”. Note that the vertical and horizontal axes are strictly t the right angle is mapped to the axes are orthogonal, and the continuous curve axis-cropping/masking and multi-curve isolation is mapped to the current cited limitation of “ processing the black pixel matrix mask to identify pixel coordinates that define the x-axis and the y-axis of the graphical KM plot”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Lahmann to have processing the black pixel matrix mask to identify pixel coordinates that define the x-axis and the y-axis of the graphical KM plot as taught by Witte in order improve the pre-processing step of image rectification by automatically tracing the grid lines and track edges of the raster log image (See Witte: Fig. 1, and [0032], “Each of the stages of the method shown in FIGS. 4A and 4B may be implemented by the system implementations shown in FIGS. 16A and 16B or by other means known to those skilled in the art. Furthermore, the processes shown in FIGS. 4A and 4B can be implemented on specialized hardware and hardware/software designed to perform the well log vectorization that is substantially more than a general computer system since the specialized hardware and hardware/software must be specially programmed or designed to be even able to perform the processes of the well log vectorization. Returning to FIGS. 4A and 4B, the method 400 may include one or more stages that are each in a processing pipeline. Although the stages are shown in a particular order in FIGS. 4A and 4B, the order of the stages may be reordered without departing from the scope of the disclosure. In a first stage, the method may load and display a raster log image (402), an example of which is shown in FIGS. 2-3. Once the raster log image is loaded into the well log vectorizer component (or the computer system), the method may perform a process of identifying and registering the raster log image components (404). For example, during this process of identifying and registering the raster log image components, the method may vertically and horizontally register the raster log image components as shown in FIGS. 5 and 6. FIG. 5 is an example of a raster well log image track that has been vertically and horizontally registered. In FIG. 5, a set of green horizontal lines (horizontal lines with periodic triangles along the line) correspond to specific measured depths (MD) of the well log such as MD=0 and MD=2000 in the example in FIG. 5. The green vertical lines (lines with squares periodically along the line) indicate the vertical edges of the track and correspond to specific measured physical values for each curve. FIG. 6 shows a raster well log image track with multiple vertical grid lines extracted (red vertical lines). Although in this image the vertical grid lines are indeed close to vertical, that is often not the case and the image can be warped to force the lines to be vertical, thus improving the image rectification prior to well log curve extraction”). Lahmann teaches a method and system that may receive a digital plot image, analyze the pixels, generate templates/series , detect axes and return digitized quantitative representations; while Witte teaches a method or system that digitize and vectorize multiple continuous curves om the raster plot image that contain axes, grids, and one or more curves/ Therefore, it is obvious for one of ordinary skill in the art to modify Lahmann by Witte to have the continuous -curve path extraction, grid/axis rectification, and background removal techniques to arrive at the claimed KM-plot digitization. The motivation to modify Lahmann by Witte is “Use of known technique to improve similar devices (methods, or products) in the same way”. However, Lahmann, modified by Witte, fails to explicitly disclose that that define the x-axis and the y-axis of the graphical KM plot and the respective digitized representation generated for each corresponding KM curve of the one or more KM curves of the graphical KM plot. However, Bekas teaches that that define the x-axis and the y-axis of the graphical KM plot and the respective digitized representation generated for each corresponding KM curve of the one or more KM curves of the graphical KM plot (See Bekas Figs. 4-5, and [0024], “Embodiments may have the beneficial effect that they are able to consider real digital images comprising graphical representations of quantitative data, e.g. scatter plots, line chart, bar chart, histogram, pie chart, flow chart or the like. The respective digital images may be generic digital images, extracted from a digital document or a scan of a printed image. The graphical representation may comprise additional structural primitives within the data region which are not representing quantitative data value but may providing additional information like a grid, a legend, or text annotations. No restrictive assumptions may be required regarding the structure of the graphical representation such as the absence of a grid or any other element in the data region that is not data. Thus, embodiments may not require to reduce the appearance variability of the graphical representation, i.e. being restricted to graphical representations with a specific predefined layout only, from which quantitative data may be extracted. Embodiments may have the beneficial effect that they for example allow automatically extracting real numerical data in original data coordinates. Thus, there may be no need for converting extracted data to real scale data manually”; [0038], “According to embodiments, the method further comprises in preparation of the extraction of the quantitative data values correcting the orientation of the graphical representation such that the first structural primitives that are labeled as axes are aligning parallel to the coordinate axes of the image coordinate system. Embodiments may have the advantage of ensuring a parallel alignment of the coordinate system of physical coordinates indicated by the axes of the graphical representation and the image coordinates indicated by the boundary of the digital image. Thus, the transformation from the image coordinates to physical coordinates of the represented quantitative data may be facilitated”; [0040], “] According to embodiments, the geometric relations between the first basic graphical objects which are provided in form of line elements comprise angles and positions of the intersections of the respective line elements. According to embodiments, the grouping of the first basic graphical objects comprises grouping the first basic graphical objects which are provided in form of characters into strings comprising one or more of the characters. Embodiments may have the advantage that structural primitives may efficiently be determined based on grouping different types of basic graphical objects differently”; [0041], “According to embodiments, the method further comprises extracting the first structural primitives which are provided in form of strings using an optical character recognition algorithm. Embodiments may have the advantage that the assignment of a semantic label may be facilitated taking into account the meaning of the structural primitives. Furthermore, structural primitives comprising only single characters or a set of characters without a literal or numerical meaning may be identified. Such structural primitives may for example markers representing quantitative data values. Also additional information about the graphical representation like a title may be extracted in order to use it for post-processing, like e.g. storing and identifying the extracted quantitative data”; [0042], “According to embodiments, the method further comprises determining the parameters of the transformation of the extracted quantitative data values from the image coordinate system to the coordinate system of physical units using the strings provided by the first structural primitives that are labeled as tick values and determining for each of the respective strings coordinate values of one or more of the first structural primitives that are labeled as a tick and associated with the string. The coordinate values are provided in units of pixels according to the image coordinate system. Embodiments may have the advantage of providing an efficient method for implementing the transformation of the extracted quantitative data values from the image coordinate system to the coordinate system of physical units in case of graphical representations comprising ticks and tick values”; and [0090], “Once the structure of a graphical representation is known, a spatial data region of the graphical representation from which quantitative data is to be extracted may be determined in block 510. In block 512, structural primitives not representing quantitative data values, like e.g. a grid, a legend and text, may be removed from inside of the axes, i.e. the data region. If necessary, the resulting graphical representation may be rotated in order to compensate for any detected skewness given by the angle at which the structural primitives labels as axes are oriented relative to the coordinate axes of the image coordinate system. Finally, only the data region is kept by cropping the image. The resulting image ideally only contains data: e.g. either markers, in the case of scatter plots, or lines with markers superimposed, in the case of line plots. In this case, the markers often represent experimental evidence, such as measurement points, and the lines depict the inferred model. In both cases the markers are of main interest”. Noe that the quantitative values are extracted from the spatial region is mapped to the cropping. And the transformation from pixel coordinate to physical unit complete the digitized data). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date that define the x-axis and the y-axis of the graphical KM plot and the respective digitized representation generated for each corresponding KM curve of the one or more KM curves of the graphical KM plot of the claimed invention was effectively filed to modify Lahmann to have processing the black pixel matrix mask to identify pixel coordinates that define the x-axis and the y-axis of the graphical KM plot as taught by Bekas in order to improve accuracy of the extracted data values (See Bekas: Fig. 1, and [0032], “The identified first structural primitives are removed from the data region. Embodiments may have the advantage that the spatial region of the graphical representation which is determined to contain the quantitative i to be extracted is cleaned up such that all what remains are structural primitives representing quantitative data. However, this cleaning up may not require any specific layout of the graphical representation, but may be applied to any arbitrary layout of graphical representation. Performing the data extraction only from a spatially restricted and cleaned up region of the graphical representation may allow an efficient data extraction independent of layout details of the graphical representation and improve the accuracy of the extracted data values”). Lahmann teaches a method and system that may generate an alternative visualization of a data set based on a specification of a selected first visualization of the data set and parameters related to the data set; while Bekas teaches a system and method that may add structural-primitive and spatial step for the axis -based cropping accuracy improvement and final mapping from pixel coordinates to real unit and no unexpected results or teaching away exist-. Therefore, it is obvious for one of ordinary skill in the art to modify Lahmann by Bekas to have determine the spatial data region framed by the axes and extract quantitative values only from the region and transform to physical units.. The motivation to modify Lahmann by Bekas is “Use of known technique to improve similar devices (methods, or products) in the same way”. However, Lahmann, modified by Witte and Bekas, fails to explicitly disclose that for each corresponding KM curve of the one or more KM curves of the graphical KM plot. However, Xiao teaches that that for each corresponding KM curve of the one or more KM curves of the graphical KM plot (See Xiao Fig. 26, and [0147], “Referring to FIG. 25, a plot 2500 of EGFR treated is shown. The EGFR TKI response prediction performance was validated in the independent dataset. Within the 87 patients who both carried sensitizing EGFR mutation and received EGFR TKI therapy, 42 patients who were predicted as responders showed significantly better OS than the 45 patients who were predicted as non-responders (P=0.024). As can be understood from the plot 2500, survival curves of predicted responders and non-responders in the EGFR TKI treated group of the validation set 116 were plotted with a Kaplan-Meier plot. All treated patients carried sensitizing EGFR mutation. The P-value was estimated with the log-rank test”). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Gibson to have define the x-axis and the y-axis of the graphical KM plot and the respective digitized representation generated for each corresponding KM curve of the one or more KM curves of the graphical KM plot of the claimed invention was effectively filed to modify Lahmann to have each corresponding KM curve of the one or more KM curves of the graphical KM plot as taught by Xiao in order to optimize patient treatment (See Xiao: Fig. 1, and [0043], “Further, in connection with clinical practice, the presently discloses technology: reduces pathologist time in analyzing pathology images to identify what may be small tumor cells, thereby expediting diagnosis and treatment; optimizes treatment for individual patients based on predicted patient prognosis; and anticipates treatment outcomes, including patient response to immunotherapy based on the spatial distribution of lymphocytes and their interaction with the tumor region, thereby further optimizing patient”). Lahmann teaches a method and system that may generate an alternative visualization of a data set based on a specification of a selected first visualization of the data set and parameters related to the data set; while Xiao teaches a system and method that may characterize patient tissue of a patient using KM estimator to process the KM plots with multiple KM curves. Therefore, it is obvious for one of ordinary skill in the art to modify Lahmann by Xiao to process the KM plot with multiple KM curves. The motivation to modify Lahmann by Xiao is “Simple substitution of one known element for another to obtain predictable results”. Regarding claim 2, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 1 as outlined above. Further, Bekas teaches that the method of claim 1, wherein the cropped colored pixel matrix mask retains only a region of the graphical KM plot that is encompassed by the positions of the x-axis and the y-axis so that the cropped colored pixel matrix exclusively contains the one or more KM curves of the graphical KM plot (See Bekas: Fig. 1, and [0034], “According to embodiments, the determining of the data region comprises identifying first structural primitives among the determined first structural primitives that are labeled as axes and determining the spatial region of the graphical representation that is framed by the axes as the data region. Embodiments may have the advantage that they provide an efficient approach to define a spatial region within which quantitative data may be found, while outside of the respective region no quantitative data may be found but rather additional information, e.g. on the physical units of the quantitative data. In case the axes are not part of a rectangle, but rather a L-shaped half rectangle, the respective half rectangle may be completed to form a full rectangle and the spatial region within the rectangle being determined as the data region”). Regarding claim 3, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 1 as outlined above. Further, Lahmann teaches that the method of claim 1, wherein the operations further comprise processing each respective group of clustered pixels to identify which respective group of clustered pixels represent a background of the graphical KM plot and which one or more respective groups of clustered pixels represent corresponding ones of the one or more KM curves of the graphical KM plot (See Lahmann: Fig. 1, and [0079], “According to embodiments, the generation of the templates comprises analyzing the identified pixel sets for identifying and filtering out pixel sets whose position, coloring, morphology and/or size indicates that said pixel set cannot represent a data point. Thereby plot labels, gridlines and/or axes, that cannot represent a single data point symbol, are filtered out. The method further comprises selectively clustering the non-filtered out pixel sets by image features into clusters of similar pixel sets. The image features are selected from a group comprising coloring features, morphological features and size. For example, all pixel sets which are red triangles may be clustered into a first cluster and all pixel sets which are black circles may be clustered into a second cluster. The method further comprises creating, selectively for each of said non-filtered out clusters, a graphical object that represents a data point symbol that is most similar to all pixel sets within said cluster and creating a template, whereby the created template comprises said graphical object as the one single data point symbol depicted in said template. For example, each feature like the color, a texture, a gradient, etc. of the graphical object represented by the cluster can be computed as the mean of the respective features of all pixel sets grouped into said cluster. The created templates may then be compared with the pixel sets for identifying completely or partially matching templates and for identifying data points at the locations in the plot where a partial or complete template match was observed”; and [0119], “According to an alternative embodiment, the assigning of the data series to the identified data points comprises assigning to each identified data point the data series represented by the template for which the data point was created. For example, the graphical object “red triangle” and the template comprising said graphical object may represent a first animal group being fed with a standard animal feed. The graphical object “black circle” and the template comprising said graphical object may represent a second animal group being fed with an improved animal food. A scatter plot may comprise pixels representing data points which indicate the size or weight of the different animal groups at a particular time or which indicate the number of animals having a particular weight or size e.g, a weight distribution plot or a size distribution plot).”). Regarding claim 4, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 1 as outlined above. Further, Lahmann teaches that the method of claim 1, wherein processing the black pixel matrix mask to identify pixel coordinates that define the x-axis and the y-axis of the graphical KM plot further comprises processing the black pixel matrix mask to identify and delineate tick mark positions along both the x-axis and the y-axis of the graphical KM plot (See: Lahmann: Fig. 1, and [0189], “Additionally, the “20” next to a tick mark on the vertical axis can be determined to match with vertical axis labels. A data point value can be determined by comparing its position to the axes label positions and interpolated values. Those values can be used to produce a new data point for each of the data points identified in the scatter plot image. This procedure can yield a dataset that includes the determined values”). Regarding claim 5, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 4 as outlined above. Further, Witte teaches that the method of claim 4, wherein generating the digitized KM plot is further based on the tick mark positions identified and delineated along both the x-axis and the y-axis of the graphical KM plot (See Witte: Figs 4A-B, and [0031], “FIGS. 4A and 4B illustrate a method 400 for well log vectorization. The method may have one or more stages. Each stage may do one or more of the following processes: 1) convert to measured physical quantities by mapping from vertical pixels to measured depth and from horizontal pixels to the well log's value; 2) store in database or write to LAS file; or 3) optionally, process the image to remove extracted curve (to facilitate extraction of other curves)”; Fig. 15, and [0057], “As shown in FIG. 15, if two interfering curves are extracted independently then both may need a large number of control points. If, after vectorizing the first curve (red curve on the left side in this example) the underlying image is modified to remove that curve, vectorization of the next curve (green on the right side in this example) may need many fewer control points. Control points for both curves are indicated as a circle with an X”. Note that the mapping step uses the scale information along the axes the tick positions information is mapped to the scale information.). Regarding claim 6, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 1 as outlined above. Further, Lahmann teaches that the method of clam 1, wherein processing the cropped colored pixel matrix to segment the colored pixels from the cropped colored pixel matrix mask into respective groups of clustered pixels (See Lahmann: Fig. 1, and [0079], “According to embodiments, the generation of the templates comprises analyzing the identified pixel sets for identifying and filtering out pixel sets whose position, coloring, morphology and/or size indicates that said pixel set cannot represent a data point. Thereby plot labels, gridlines and/or axes, that cannot represent a single data point symbol, are filtered out. The method further comprises selectively clustering the non-filtered out pixel sets by image features into clusters of similar pixel sets. The image features are selected from a group comprising coloring features, morphological features and size. For example, all pixel sets which are red triangles may be clustered into a first cluster and all pixel sets which are black circles may be clustered into a second cluster. The method further comprises creating, selectively for each of said non-filtered out clusters, a graphical object that represents a data point symbol that is most similar to all pixel sets within said cluster and creating a template, whereby the created template comprises said graphical object as the one single data point symbol depicted in said template. For example, each feature like the color, a texture, a gradient, etc. of the graphical object represented by the cluster can be computed as the mean of the respective features of all pixel sets grouped into said cluster. The created templates may then be compared with the pixel sets for identifying completely or partially matching templates and for identifying data points at the locations in the plot where a partial or complete template match was observed”. Note that selectively clustering the non-filtered out pixel sets by image features into clusters of similar pixel sets is mapped to processing the cropped colored pixel matrix to segment the colored pixels from the cropped colored pixel matrix mask into respective groups of clustered pixels) comprises: flattening the cropped colored pixel matrix mask into a vector of color code vectors (See Lahmann: Fig. 1, and [0023], “The pixel sets are identified in the received scatter plot image or in a derivative thereof for identifying graphical objects in the scatter plot image. Each identified pixel set is assumed to represent a respective graphical object in the scatter plot image. The received digital image may have multiple forms, e.g. a binary (“black and white”) image, a single-channel (“graylever”) image or a multi-channel (e.g. RGB or (MYK) image. Optionally, a multi-channel image may be transformed into one or more single-channel images and/or the one or more single-channel image may be transformed into a binary image in additional processing steps that are performed for preparing the image data for identifying the graphical objects in the form of pixel sets in the received digital image” Note that the color based clustering is standard routing preprocessing to re[resent the color in RGB code values.); and processing the vector of color code vectors using a K-means clustering for clustering the respective colors associated with the one or more KM curves into the respective groups of clustered pixels (See Lahmann: Fig. 1, and [0016], “The accuracy of identifying individual data points in a plot and of identifying the data series a data point belongs to may be greatly increased. In particular, the method may be much more robust against data point detection errors and data series assignment errors which may result from an overlap of two or more data points of the same or of different data series. For example in case a clustering algorithm is applied on the plot for identifying the data points and their respective series in a single clustering step, the problem arises that overlapping data points may be erroneously identified as a new type of data point symbol and as a new type of data series. This may be prevented by identifying templates which respectively comprise a (single) data point symbol (of the data series represented by said template), and then using said templates in a further, separate step for identifying the actual data points (data series instances)”. Note that K-means is he most commonly used in pixel clustering algorithm). Regarding claim 7, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 6 as outlined above. Further, Lahmann teaches that the method of claim 6, wherein each respective group of clustered pixels has a centroid defining the different respective color associated with the pixels in the respective group of clustered pixels (See Lahmann: Fig. 1, and [0113], “After the function finishes the comparison, the best matches can be found as local minimums (when “sum of squared differences” was used) or maximums (when “correlation coefficient” or “cross correlation” was used). In case of a color image, template summation in the numerator and each sum in the denominator is done over all of the channels and separate mean values are used for each channel. Alternatively, sum of squared differences may be calculated as the sum of the squares of the norm of the difference between the color intensity vectors of a multi-channel template and the patch image. That is, the function can take a color template and a color image. The result is preferably a single-channel image, which is easier to analyze”. Note that the mean values is mapped to the centroid). Regarding claim 11, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 1 as outlined above. Further, Xiao teaches that the method of claim 1, wherein the operations further comprise executing an independent patient data (IPD) extraction process that uses number at risk data obtained from the graphical KM plot to extract IPD from the digitized KM plot (See Xiao: Fig. 1, and [0091], “In one implementation, the system 100 predicts treatment response to EGFR TKI targeted therapy using a deep learning pathology image analysis pipeline. The cell organization-based GCN of the GCN system 120 generates a prognosis for a patient by providing all connected graphs from the pathology image 104 in a single disconnected graph. The disconnected graph is input at the input layer into a plurality of EC convolutional layers and output into a global mean-pooling layer. The global mean pooling layer evaluates all cell types within the TME. The output from the global mean pooling layer is input into the softmax layer, which outputs an output layer providing a probability of high-risk for the prognostic model 108. Nuclei morphological features may be set to 1 to be excluded from the input graph in generating a response prediction using the GCN system 120”; [0093], “To train response prediction of the GCN system 120 in this example, cross-entropy may be used as a loss function and an adaptive deep learning rate with scaling factor of 2 may be used as the optimizer. A maximum training epoch is set as 300, and the model at the 135.sup.th epoch with a highest classification accuracy in the training set 114 is selected. The probability of belonging to the benefitting group is used as a benefitting score. In the testing set 114, patients are dichotomized into the benefitting and non-benefitting groups according to the median benefitting score. Kaplan-Meier curves and log-rank tests may illustrate the survival difference between EGFR TKI treated and non-treated patients in the benefitting group and non-benefitting group, respectively. The differences are considered significant when two-tailed p-value<0.05”; and [0100], “The system 100 provides GCN-based pathology image analysis for histological classification in a variety of contexts. As described herein, in one example, the system 100 predicts responsiveness to EGFR TKI targeted therapy. Turning to FIG. 14, in one example, to predict a slide-level benefitting score, all image patches from the same pathology slide are grouped together to construct one disconnected graph. The GCN system 120 is trained to predict a benefitting score for each input graph and applied to patients with EGFR mutation in the testing set 118. As shown in a plot 1400 of survival proportion and time after metastasis, within the predicted benefitting group, patients who did not receive EGFR TKI targeted therapy showed significantly worse survival than patients who received EGFR TKI targeted therapy. As shown in the plot 1400, p=0.0002; without targeted therapy versus with targeted therapy, with a Hazard Ratio [HR]=6.81, and a 95% Confidence Interval [CI] 2.14-21.73. In contrast, within the predicted non-benefitting group, there was no significant survival difference between targeted therapy treated and non-treated patient groups. As shown in the plot 1400, p=0.10; without targeted therapy versus with targeted therapy, HR=2.22, and 95% CI 0.85-5.83. As further illustrated in Table 2, after adjusting for potential clinical confounders, including age, gender, smoking status, surgery, and stage at diagnosis, a high benefitting score calculated by the GCN system 120 is predictive for prolonged overall survival in patients who carried the EGFR mutation and received EGFR TKI targeted therapy”). Regarding claim 12, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 11 as outlined above. Further, Xiao teaches that the method of claim 11, wherein the operations further comprise generating a digitized IPD plot that conveys the IPD extracted from the digitized KM plot (See Xiao: Fig. 1, and [0107], “Although some deep-learning algorithms perform well in many preset computational challenges in the biomedical space, the performances of such algorithms often significantly deteriorate when applied to digital pathology images, due to the diversity of pathological images. For example, the differences between scanner type, manufacturer and digitization process can restrain the quality of digital pathology slides in terms of image sharpness, resolution, noise level, amplification magnitude, and/or the like. In the context of H&E-stained histopathology images analysis described herein, the color variation provides a unique challenge, which can be impacted by stain concentration, time elapsed, environmental temperatures upon staining, and/or other factors. Without properly accounting for these image quality issues and staining variations, classification, segmentation, and characterization may have decreased accuracy. In other words, inadequate image quality, low amplification magnitude, and staining variation may decrease an accuracy of tumor region segmentation, nuclei detection, and classification. Thus, the system 100 may perform image restoration and quality enhancement using the image restoration system 122 to restore blurred regions, enhance low resolution/magnification into high resolution, normalize staining colors to reduce staining variation, and/or the like”). Regarding claim 13, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 1 as outlined above. Further, Xiao teaches that the method of claim 1, wherein each KM curve of the one or more KM curves of the graphical KM plot depicts survival probability over time for a respective group of subjects (See Xiao: Fig. 14 and [0100], “The system 100 provides GCN-based pathology image analysis for histological classification in a variety of contexts. As described herein, in one example, the system 100 predicts responsiveness to EGFR TKI targeted therapy. Turning to FIG. 14, in one example, to predict a slide-level benefitting score, all image patches from the same pathology slide are grouped together to construct one disconnected graph. The GCN system 120 is trained to predict a benefitting score for each input graph and applied to patients with EGFR mutation in the testing set 118. As shown in a plot 1400 of survival proportion and time after metastasis, within the predicted benefitting group, patients who did not receive EGFR TKI targeted therapy showed significantly worse survival than patients who received EGFR TKI targeted therapy. As shown in the plot 1400, p=0.0002; without targeted therapy versus with targeted therapy, with a Hazard Ratio [HR]=6.81, and a 95% Confidence Interval [CI] 2.14-21.73. In contrast, within the predicted non-benefitting group, there was no significant survival difference between targeted therapy treated and non-treated patient groups. As shown in the plot 1400, p=0.10; without targeted therapy versus with targeted therapy, HR=2.22, and 95% CI 0.85-5.83. As further illustrated in Table 2, after adjusting for potential clinical confounders, including age, gender, smoking status, surgery, and stage at diagnosis, a high benefitting score calculated by the GCN system 120 is predictive for prolonged overall survival in patients who carried the EGFR mutation and received EGFR TKI targeted therapy”). Regarding claim 14, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 1 as outlined above. Further, Lahmann, Witte, Bekas and Xiao teach that a system (See Lahmann: Fig. 2, and [0168], “FIG. 2 depicts a computer system 200 of an embodiment of the invention that is configured for extracting data from a scatter plot 218. The extraction of data from a scatter plot may be advantageous as scatter plots are commonly used. The extraction of data from scatter plots has often been reported to be difficult due to problems to correctly identify the data points and series to which a particular data point belongs. Embodiments of the invention may allow the extraction of data from scatter plots in an efficient, error robust and accurate manner”) comprising: data processing hardware(See Lahmann: Fig. 2, and [0126], “In a further aspect, the invention relates to a tangible non-volatile storage medium comprising computer-interpretable instructions stored thereon. The instructions, when executed by a processor, cause the processor to perform a method for extracting data from a scatter plot. The method comprises”); and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations (See Lahmann: Fig. 2, and [0096], ““Creating a data point” in this context may mean that pixels in the digital plot image are identified to represent a data point of a particular data series and a corresponding data object, e.g. a class instance or a data array is created and stored in the main memory and optionally also in a non-volatile storage medium. These created data points may then be stored in any data format of interest, e.g. as a table, as a comma delimited file, as a database record, or as a class object instance of an application program written in an object oriented programming language”) comprising: receiving an input image containing a graphical Kaplan-Meier (KM) plot of one or more KM curves, the graphical KM plot having an x-axis and a y-axis orthogonal to the x-axis(See Lahmann: Fig. 1, and [0160], “In a first step 102, the image analysis logic receives a digital image of a scatter plot. For example, the image analysis logic can read a JPEG RGB image that depicts a scatter plot from a storage medium or from a webpage that comprises the JPEG image. Alternatively, the image analysis logic can be coupled to an image acquisition system, e.g. a camera or a scanner, and receive the image from the image acquisition system. Then, the received image can optionally be processed for transforming the RGB image into a binary digital image. Alternatively, the digital image of the scatter plot can already be received in the form of a binary scatter plot imager”) ; processing the input image to convert the graphical KM plot into a three-dimensional (3D) array representing a color of each pixel from the graphical KM plot by a respective 3D vector ; (See Lahmann: Figs. 4A-D, and [0181], “In a step depicted in FIG. 4A, adjacent pixel groups of pixels which are similar to each other are identified, e.g. by means of a connected-component analysis, as “pixel sets” 402. For example, the pixels of the axes or of one axes may correspond to one pixel set. The pixels of each character of an axis label or of the title may correspond to one pixel set. The pixels of each symbol representing an isolated data point may correspond to a respective single pixel set. The pixels of each cluster of overlapping symbols of multiple data points may correspond to a respective single pixel set. Thus, the pixel sets or “blobs” depicted in FIG. 4A may represent data points, sets of overlapping data points, artifacts, labels, axes, characters and other object types that might cause errors”; [0203], “An inner morphological gradient filter, performed by taking a morphologically dilated image minus the original image, is applied 34 to each red-green-blue (RGB) color channel of the original image to produce three new single-channel (grayscale) images, a, b, and c. A composite grayscale image, x, is then computed from a, b, and c by selecting the maximum pixel value at each pixel coordinate in an image from images a, b, and c and storing the selected maximum pixel value into x. One or more optimal threshold computations (such as the global statistical mean and the standard deviation of pixel intensities within x), are performed 36 on x to produce a binary image featuring the contours of the graphical objects in the original image. The collection of individual connected components in the image is computed and each of these elements is used to segment, locate, and label the set of data components in the image as outlined below”. Note that the color RGB is mapped to the 3D vector per pixel, and the digital color image pixels are mapped to arrays) processing the 3D array to generate (See Lahmann: Fig. 4A-D, and [0024], “In case one or more derivative images, e.g. a binary image, is created, the template matching can be performed on a derivative image (which is typically faster) or on the original image (which is typically more accurate as it may comprise more information than a derivative image). Depending on the form of the received image and other factors (e.g. whether high accuracy or high performance is preferred), different kinds of pixel set detection algorithms, template matching approaches and/or data series assignment approaches can be applied”): a black pixel matrix mask by converting all white pixels to pure white and all other pixels to pure black (See Lahmann: Figs. 4A-D, and [0030], “The identification of pixel sets based on similar adjacent pixels can be performed e.g. by means of a connected component analysis which may be performed e.g. on an original multi-channel scatter plot image. For instance, the connected component analysis may be performed on an RGB scatter plot image. Alternatively, the identification of pixel sets based on similar adjacent pixels can be performed on a derivative image, e.g. a grayscale or binary black-white image”; and [0048], “According to embodiments, binary (black-and-white) images are generated as an intermediate step to the extraction of the connected components, i.e. pixel sets, for reducing the color information. Then, all white, repetitive black (depending on the definition) “pixel blobs” are used as “masks” which identify respective pixel sets in the multi-channel image which are further processed for extracting the templates and for performing the template mapping”); and a colored pixel matrix mask by converting all black, grey, and white pixels to pure white (See Lahmann: Figs. 4A-D, [0048], “According to embodiments, binary (black-and-white) images are generated as an intermediate step to the extraction of the connected components, i.e. pixel sets, for reducing the color information. Then, all white, repetitive black (depending on the definition) “pixel blobs” are used as “masks” which identify respective pixel sets in the multi-channel image which are further processed for extracting the templates and for performing the template mapping””); processing the black pixel matrix mask to identify pixel coordinates (See Lahmann: Fig. 7, and [0203], “An inner morphological gradient filter, performed by taking a morphologically dilated image minus the original image, is applied 34 to each red-green-blue (RGB) color channel of the original image to produce three new single-channel (grayscale) images, a, b, and c. A composite grayscale image, x, is then computed from a, b, and c by selecting the maximum pixel value at each pixel coordinate in an image from images a, b, and c and storing the selected maximum pixel value into x. One or more optimal threshold computations (such as the global statistical mean and the standard deviation of pixel intensities within x), are performed 36 on x to produce a binary image featuring the contours of the graphical objects in the original image. The collection of individual connected components in the image is computed and each of these elements is used to segment, locate, and label the set of data components in the image as outlined below”; and [0207], “Once all data points in the plot image are identified, a bounding box is initialized 44 enclosing all elements of the data set. In some embodiments, the bounding box is deformed so that its edges reside on the vertical and horizontal axes lines of the plot image. The axes are then identified 46 as the line segments representing the edges of the bounding box. The textual labels, including series labels, chart title(s), axes range values, etc. are identified and extracted using optical character recognition (OCR)”) that define the x-axis and the y-axis of the graphical KM plot (See Witte Figs. 4-5, and [0034], “The processing of the image may include a number of subprocesses. In particular, the stage may include: 1) detecting and removing horizontal and vertical grid lines from the image in a known manner; 2) optionally removing isolated noise by detecting and removing connected components below some number of connected pixels; 3) optionally smoothing the image in the horizontal direction using, for example, an exponential or Gaussian smoothing function; 4) optionally extracting measures of local image orientation and structure using, for example, the eigenvalues and eigenvectors of the local gradient structure tensor at each point in the image; and 5) optionally enhancing the curves in the image by, for example, simple color separation if each curve is a different color or, for example, by performing a local Hough transform or anisotropic diffusion at each point in a binary or gray-scale image”; Figs. 2-3and [0030], “Prior to describing the digital well log vectorization process, a raster well log image and the nomenclature for the regions in the raster well log image are described with reference to FIG. 1. As shown, the raster well log image may have a Header region, a Tool Configuration region, a Miscellaneous region and a Trailer region that each contain meta-information about the well, well bore, measurement devices and location. The digital well log vectorization process described below may be used for processing of the Upper Sections (US), Lower Section (LS) and the Log Sections. The Upper and Lower Sections contain scale information for the well log measurements. The Log Section contains the raster image of the well log curves. In the other figures described below, the Log sections only are shown. For example, FIGS. 2 and 3 illustrate two examples of a raster well log image containing two tracks and multiple well log curves. A left hand side of FIG. 2 is an example of a raster well log image containing two tracks and multiple well log curves. Note that in FIG. 2, the ‘horizontal’ and ‘vertical’ grid lines are neither horizontal nor vertical, indicating rotation and possible additional horizontal skew or stretch of the image. Also note the variable image quality, including a dark splotch in the lower part of the image. The right hand side of FIG. 2 shows a portion of the image in the left hand side of FIG. 2 after well-known warping the image to rectify the grid lines. FIG. 3 shows another raster well log image with two tracks, each containing multiple curves. Note the vertical scale between the tracks. This scale is used to convert from vertical pixels to measured depth along the well bore”; and fig. 6, and [0027]. “A further advantage of the system and method may be that the image-guided shortest path algorithm can also be used to improve the pre-processing step of image rectification, by automatically tracing the grid lines and track edges of the raster log image. These extracted grid lines should be strictly horizontal or vertical and intersect at right angles. Detected deviations from these conditions form the basis for accurately warping the image to remove distortions due to poor scanning or damage to the original paper records. Improved image rectification results in more accurate values for the extracted digital well log curves”; [0032], “Each of the stages of the method shown in FIGS. 4A and 4B may be implemented by the system implementations shown in FIGS. 16A and 16B or by other means known to those skilled in the art. Furthermore, the processes shown in FIGS. 4A and 4B can be implemented on specialized hardware and hardware/software designed to perform the well log vectorization that is substantially more than a general computer system since the specialized hardware and hardware/software must be specially programmed or designed to be even able to perform the processes of the well log vectorization. Returning to FIGS. 4A and 4B, the method 400 may include one or more stages that are each in a processing pipeline. Although the stages are shown in a particular order in FIGS. 4A and 4B, the order of the stages may be reordered without departing from the scope of the disclosure. In a first stage, the method may load and display a raster log image (402), an example of which is shown in FIGS. 2-3. Once the raster log image is loaded into the well log vectorizer component (or the computer system), the method may perform a process of identifying and registering the raster log image components (404). For example, during this process of identifying and registering the raster log image components, the method may vertically and horizontally register the raster log image components as shown in FIGS. 5 and 6. FIG. 5 is an example of a raster well log image track that has been vertically and horizontally registered. In FIG. 5, a set of green horizontal lines (horizontal lines with periodic triangles along the line) correspond to specific measured depths (MD) of the well log such as MD=0 and MD=2000 in the example in FIG. 5. The green vertical lines (lines with squares periodically along the line) indicate the vertical edges of the track and correspond to specific measured physical values for each curve. FIG. 6 shows a raster well log image track with multiple vertical grid lines extracted (red vertical lines). Although in this image the vertical grid lines are indeed close to vertical, that is often not the case and the image can be warped to force the lines to be vertical, thus improving the image rectification prior to well log curve extraction”. Note that the vertical and horizontal axes are strictly t the right angle is mapped to the axes are orthogonal, and the continuous curve axis-cropping/masking and multi-curve isolation is mapped to the current cited limitation of “ processing the black pixel matrix mask to identify pixel coordinates that define the x-axis and the y-axis of the graphical KM plot”); cropping the colored pixel matrix mask based on the identified pixel coordinates that define the x-axis and the y-axis of the graphical KM plot (See Lahmann: Fig. 7, and [0207], “Once all data points in the plot image are identified, a bounding box is initialized 44 enclosing all elements of the data set. In some embodiments, the bounding box is deformed so that its edges reside on the vertical and horizontal axes lines of the plot image. The axes are then identified 46 as the line segments representing the edges of the bounding box. The textual labels, including series labels, chart title(s), axes range values, etc. are identified and extracted using optical character recognition (OCR)”); processing the cropped colored pixel matrix mask to segment the colored pixels from the cropped colored pixel matrix mask into respective groups of clustered pixels, each respective group of clustered pixels associated with a different respective color (See Lahmann: Fig. 4A-D, and [0137], comparing each of the templates with pixels of a target image, the target image being the received scatter plot image or the derivative of the received scatter plot image, for identifying positions of matching templates, a matching template being a template whose degree of similarity to pixels of the target image exceeds a similarity threshold”; and [0162], “In step 106, the image analysis logic analyzes the identified pixel sets in order to generate a plurality of templates. The template generation may involve the generation of template candidates from which the finally used templates are selected in one or more filtering steps as described, for example, with reference to FIG. 4. Each template is a pixel structure that depicts exactly one data point symbol, e.g. a red triangle or a black circle. Thus, each template and respective data point symbol represents a respective data series, whereby all data points of a particular data series are assumed to be represented in the plot with the respective symbol of that series”. Note that the plurality of templates is mapped to the respective groups of clustered pixels, each respective group of clustered pixels associated with a different respective color); processing each respective group of clustered pixels that represents a corresponding KM curve of the one or more KM curves of the graphical KM plot to generate a respective digitized representation of the corresponding KM curve (See Lahmann: Fig. 4A-D and 7, and [0206], “The identified set of data points are segmented 42 into different data series each including a plurality of data points, based, for instance, on the locations, spacings, coloring, patterns, and/or shapes of the image elements they represent. For instance, if a line chart plot image contained three lines of different colors, red, blue, and yellow, the digitization system segments the data into three separate series, with data sets corresponding to each line based on color. Similarly, if a scatter plot image contained two types of data point elements, circles and diamonds, the digitization system segments the data into two separate series, with one data set corresponding to all circle elements of the plot image and one data set corresponding to all diamond elements of the plot image. As discussed below, the different series are identified with distinct markers and are separated into partially or wholly distinct data sets in the data grid”, Note that the segment data point in series is the digitized representation of that curve/series, and it is mapped to the KM curves); and generating a digitized KM plot based on the identified pixel coordinates (See Lahmann: Fig. 7, and[0014], “returning the identified data points and the data series to which it is assigned”; [0208], “All extracted components of the plot image, including data points, axes markers, and textual labels (series labels, chart title(s), axes range values, etc.) are visually presented in the user interface 50, as specified below in steps 52-56”; and [0209], “The identified data points are marked 52, and the identified axes are marked 54 with polygons, crosshairs, or lines on the canvas overlaying the plot image. The textual labels, including series labels, chart title(s), axes range values, etc., are visually presented in the user interface 56 such that the user may manipulate these elements”. Note that the returned identified data points are mapped to the digitized KM plot) that define the x-axis and the y-axis of the graphical KM plot and the respective digitized representation generated for each corresponding KM curve (See Bekas Figs. 4-5, and [0024], “Embodiments may have the beneficial effect that they are able to consider real digital images comprising graphical representations of quantitative data, e.g. scatter plots, line chart, bar chart, histogram, pie chart, flow chart or the like. The respective digital images may be generic digital images, extracted from a digital document or a scan of a printed image. The graphical representation may comprise additional structural primitives within the data region which are not representing quantitative data value but may providing additional information like a grid, a legend, or text annotations. No restrictive assumptions may be required regarding the structure of the graphical representation such as the absence of a grid or any other element in the data region that is not data. Thus, embodiments may not require to reduce the appearance variability of the graphical representation, i.e. being restricted to graphical representations with a specific predefined layout only, from which quantitative data may be extracted. Embodiments may have the beneficial effect that they for example allow automatically extracting real numerical data in original data coordinates. Thus, there may be no need for converting extracted data to real scale data manually”; [0038], “According to embodiments, the method further comprises in preparation of the extraction of the quantitative data values correcting the orientation of the graphical representation such that the first structural primitives that are labeled as axes are aligning parallel to the coordinate axes of the image coordinate system. Embodiments may have the advantage of ensuring a parallel alignment of the coordinate system of physical coordinates indicated by the axes of the graphical representation and the image coordinates indicated by the boundary of the digital image. Thus, the transformation from the image coordinates to physical coordinates of the represented quantitative data may be facilitated”; [0040], “] According to embodiments, the geometric relations between the first basic graphical objects which are provided in form of line elements comprise angles and positions of the intersections of the respective line elements. According to embodiments, the grouping of the first basic graphical objects comprises grouping the first basic graphical objects which are provided in form of characters into strings comprising one or more of the characters. Embodiments may have the advantage that structural primitives may efficiently be determined based on grouping different types of basic graphical objects differently”; [0041], “According to embodiments, the method further comprises extracting the first structural primitives which are provided in form of strings using an optical character recognition algorithm. Embodiments may have the advantage that the assignment of a semantic label may be facilitated taking into account the meaning of the structural primitives. Furthermore, structural primitives comprising only single characters or a set of characters without a literal or numerical meaning may be identified. Such structural primitives may for example markers representing quantitative data values. Also additional information about the graphical representation like a title may be extracted in order to use it for post-processing, like e.g. storing and identifying the extracted quantitative data”; [0042], “According to embodiments, the method further comprises determining the parameters of the transformation of the extracted quantitative data values from the image coordinate system to the coordinate system of physical units using the strings provided by the first structural primitives that are labeled as tick values and determining for each of the respective strings coordinate values of one or more of the first structural primitives that are labeled as a tick and associated with the string. The coordinate values are provided in units of pixels according to the image coordinate system. Embodiments may have the advantage of providing an efficient method for implementing the transformation of the extracted quantitative data values from the image coordinate system to the coordinate system of physical units in case of graphical representations comprising ticks and tick values”; and [0090], “Once the structure of a graphical representation is known, a spatial data region of the graphical representation from which quantitative data is to be extracted may be determined in block 510. In block 512, structural primitives not representing quantitative data values, like e.g. a grid, a legend and text, may be removed from inside of the axes, i.e. the data region. If necessary, the resulting graphical representation may be rotated in order to compensate for any detected skewness given by the angle at which the structural primitives labels as axes are oriented relative to the coordinate axes of the image coordinate system. Finally, only the data region is kept by cropping the image. The resulting image ideally only contains data: e.g. either markers, in the case of scatter plots, or lines with markers superimposed, in the case of line plots. In this case, the markers often represent experimental evidence, such as measurement points, and the lines depict the inferred model. In both cases the markers are of main interest”. Noe that the quantitative values are extracted from the spatial region is mapped to the cropping. And the transformation from pixel coordinate to physical unit complete the digitized data) of the one or more KM curves of the graphical KM plot (See Xiao Fig. 26, and [0147], “Referring to FIG. 25, a plot 2500 of EGFR treated is shown. The EGFR TKI response prediction performance was validated in the independent dataset. Within the 87 patients who both carried sensitizing EGFR mutation and received EGFR TKI therapy, 42 patients who were predicted as responders showed significantly better OS than the 45 patients who were predicted as non-responders (P=0.024). As can be understood from the plot 2500, survival curves of predicted responders and non-responders in the EGFR TKI treated group of the validation set 116 were plotted with a Kaplan-Meier plot. All treated patients carried sensitizing EGFR mutation. The P-value was estimated with the log-rank test”). Regarding claim 15, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 14 as outlined above. Further, Bekas teaches that the system of claim 14, wherein the cropped colored pixel matrix mask retains only a region of the graphical KM plot that is encompassed by the positions of the x-axis and the y-axis so that the cropped colored pixel matrix exclusively contains the one or more KM curves of the graphical KM plot (See Bekas: Fig. 1, and [0034], “According to embodiments, the determining of the data region comprises identifying first structural primitives among the determined first structural primitives that are labeled as axes and determining the spatial region of the graphical representation that is framed by the axes as the data region. Embodiments may have the advantage that they provide an efficient approach to define a spatial region within which quantitative data may be found, while outside of the respective region no quantitative data may be found but rather additional information, e.g. on the physical units of the quantitative data. In case the axes are not part of a rectangle, but rather a L-shaped half rectangle, the respective half rectangle may be completed to form a full rectangle and the spatial region within the rectangle being determined as the data region”). Regarding claim 16, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 14 as outlined above. Further, Lahmann teaches that the system of claim 14, wherein the operations further comprise processing each respective group of clustered pixels to identify which respective group of clustered pixels represent a background of the graphical KM plot and which one or more respective groups of clustered pixels represent corresponding ones of the one or more KM curves of the graphical KM plot (See Lahmann: Fig. 1, and [0079], “According to embodiments, the generation of the templates comprises analyzing the identified pixel sets for identifying and filtering out pixel sets whose position, coloring, morphology and/or size indicates that said pixel set cannot represent a data point. Thereby plot labels, gridlines and/or axes, that cannot represent a single data point symbol, are filtered out. The method further comprises selectively clustering the non-filtered out pixel sets by image features into clusters of similar pixel sets. The image features are selected from a group comprising coloring features, morphological features and size. For example, all pixel sets which are red triangles may be clustered into a first cluster and all pixel sets which are black circles may be clustered into a second cluster. The method further comprises creating, selectively for each of said non-filtered out clusters, a graphical object that represents a data point symbol that is most similar to all pixel sets within said cluster and creating a template, whereby the created template comprises said graphical object as the one single data point symbol depicted in said template. For example, each feature like the color, a texture, a gradient, etc. of the graphical object represented by the cluster can be computed as the mean of the respective features of all pixel sets grouped into said cluster. The created templates may then be compared with the pixel sets for identifying completely or partially matching templates and for identifying data points at the locations in the plot where a partial or complete template match was observed”; and [0119], “According to an alternative embodiment, the assigning of the data series to the identified data points comprises assigning to each identified data point the data series represented by the template for which the data point was created. For example, the graphical object “red triangle” and the template comprising said graphical object may represent a first animal group being fed with a standard animal feed. The graphical object “black circle” and the template comprising said graphical object may represent a second animal group being fed with an improved animal food. A scatter plot may comprise pixels representing data points which indicate the size or weight of the different animal groups at a particular time or which indicate the number of animals having a particular weight or size e.g, a weight distribution plot or a size distribution plot)”). Regarding claim 17, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 14 as outlined above. Further, Lahmann teaches that the system of claim 14, wherein processing the black pixel matrix mask to identify pixel coordinates that define the x-axis and the y-axis of the graphical KM plot further comprises processing the black pixel matrix mask to identify and delineate tick mark positions along both the x-axis and the y-axis of the graphical KM plot (See: Lahmann: Fig. 1, and [0189], “Additionally, the “20” next to a tick mark on the vertical axis can be determined to match with vertical axis labels. A data point value can be determined by comparing its position to the axes label positions and interpolated values. Those values can be used to produce a new data point for each of the data points identified in the scatter plot image. This procedure can yield a dataset that includes the determined values”). Regarding claim 18, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 17 as outlined above. Further, Witte teaches that the system of claim 17, wherein generating the digitized KM plot is further based on the tick mark positions identified and delineated along both the x-axis and the y-axis of the graphical KM plot (See Witte: Figs 4A-B, and [0031], “FIGS. 4A and 4B illustrate a method 400 for well log vectorization. The method may have one or more stages. Each stage may do one or more of the following processes: 1) convert to measured physical quantities by mapping from vertical pixels to measured depth and from horizontal pixels to the well log's value; 2) store in database or write to LAS file; or 3) optionally, process the image to remove extracted curve (to facilitate extraction of other curves)”; Fig. 15, and [0057], “As shown in FIG. 15, if two interfering curves are extracted independently then both may need a large number of control points. If, after vectorizing the first curve (red curve on the left side in this example) the underlying image is modified to remove that curve, vectorization of the next curve (green on the right side in this example) may need many fewer control points. Control points for both curves are indicated as a circle with an X”. Note that the mapping step uses the scale information along the axes the tick positions information is mapped to the scale information.). Regarding claim 19, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 14 as outlined above. Further, Lahmann teaches that the system of claim 14, wherein processing the cropped colored pixel matrix to segment the colored pixels from the cropped colored pixel matrix mask into respective groups of clustered pixels comprises: flattening the cropped colored pixel matrix mask into a vector of color code vectors (See Lahmann: Fig. 1, and [0023], “The pixel sets are identified in the received scatter plot image or in a derivative thereof for identifying graphical objects in the scatter plot image. Each identified pixel set is assumed to represent a respective graphical object in the scatter plot image. The received digital image may have multiple forms, e.g. a binary (“black and white”) image, a single-channel (“graylever”) image or a multi-channel (e.g. RGB or (MYK) image. Optionally, a multi-channel image may be transformed into one or more single-channel images and/or the one or more single-channel image may be transformed into a binary image in additional processing steps that are performed for preparing the image data for identifying the graphical objects in the form of pixel sets in the received digital image” Note that the color based clustering is standard routing preprocessing to re[resent the color in RGB code values.); and processing the vector of color code vectors using a K-means clustering for clustering the respective colors associated with the one or more KM curves into the respective groups of clustered pixels (See Lahmann: Fig. 1, and [0016], “The accuracy of identifying individual data points in a plot and of identifying the data series a data point belongs to may be greatly increased. In particular, the method may be much more robust against data point detection errors and data series assignment errors which may result from an overlap of two or more data points of the same or of different data series. For example in case a clustering algorithm is applied on the plot for identifying the data points and their respective series in a single clustering step, the problem arises that overlapping data points may be erroneously identified as a new type of data point symbol and as a new type of data series. This may be prevented by identifying templates which respectively comprise a (single) data point symbol (of the data series represented by said template), and then using said templates in a further, separate step for identifying the actual data points (data series instances)”. Note that K-means is he most commonly used in pixel clustering algorithm). Regarding claim 20, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 19 as outlined above. Further, Lahmann teaches that the system of claim 19, wherein each respective group of clustered pixels has a centroid defining the different respective color associated with the pixels in the respective group of clustered pixels (See Lahmann: Fig. 1, and [0113], “After the function finishes the comparison, the best matches can be found as local minimums (when “sum of squared differences” was used) or maximums (when “correlation coefficient” or “cross correlation” was used). In case of a color image, template summation in the numerator and each sum in the denominator is done over all of the channels and separate mean values are used for each channel. Alternatively, sum of squared differences may be calculated as the sum of the squares of the norm of the difference between the color intensity vectors of a multi-channel template and the patch image. That is, the function can take a color template and a color image. The result is preferably a single-channel image, which is easier to analyze”. Note that the mean values is mapped to the centroid). Regarding claim 24, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 14 as outlined above. Further, Xiao teaches that the system of claim 14, wherein the operations further comprise executing an independent patient data (IPD) extraction process that uses number at risk data obtained from the graphical KM plot to extract IPD from the digitized KM plot (See Xiao: Fig. 1, and [0091], “In one implementation, the system 100 predicts treatment response to EGFR TKI targeted therapy using a deep learning pathology image analysis pipeline. The cell organization-based GCN of the GCN system 120 generates a prognosis for a patient by providing all connected graphs from the pathology image 104 in a single disconnected graph. The disconnected graph is input at the input layer into a plurality of EC convolutional layers and output into a global mean-pooling layer. The global mean pooling layer evaluates all cell types within the TME. The output from the global mean pooling layer is input into the softmax layer, which outputs an output layer providing a probability of high-risk for the prognostic model 108. Nuclei morphological features may be set to 1 to be excluded from the input graph in generating a response prediction using the GCN system 120”; [0093], “To train response prediction of the GCN system 120 in this example, cross-entropy may be used as a loss function and an adaptive deep learning rate with scaling factor of 2 may be used as the optimizer. A maximum training epoch is set as 300, and the model at the 135.sup.th epoch with a highest classification accuracy in the training set 114 is selected. The probability of belonging to the benefitting group is used as a benefitting score. In the testing set 114, patients are dichotomized into the benefitting and non-benefitting groups according to the median benefitting score. Kaplan-Meier curves and log-rank tests may illustrate the survival difference between EGFR TKI treated and non-treated patients in the benefitting group and non-benefitting group, respectively. The differences are considered significant when two-tailed p-value<0.05”; and [0100], “The system 100 provides GCN-based pathology image analysis for histological classification in a variety of contexts. As described herein, in one example, the system 100 predicts responsiveness to EGFR TKI targeted therapy. Turning to FIG. 14, in one example, to predict a slide-level benefitting score, all image patches from the same pathology slide are grouped together to construct one disconnected graph. The GCN system 120 is trained to predict a benefitting score for each input graph and applied to patients with EGFR mutation in the testing set 118. As shown in a plot 1400 of survival proportion and time after metastasis, within the predicted benefitting group, patients who did not receive EGFR TKI targeted therapy showed significantly worse survival than patients who received EGFR TKI targeted therapy. As shown in the plot 1400, p=0.0002; without targeted therapy versus with targeted therapy, with a Hazard Ratio [HR]=6.81, and a 95% Confidence Interval [CI] 2.14-21.73. In contrast, within the predicted non-benefitting group, there was no significant survival difference between targeted therapy treated and non-treated patient groups. As shown in the plot 1400, p=0.10; without targeted therapy versus with targeted therapy, HR=2.22, and 95% CI 0.85-5.83. As further illustrated in Table 2, after adjusting for potential clinical confounders, including age, gender, smoking status, surgery, and stage at diagnosis, a high benefitting score calculated by the GCN system 120 is predictive for prolonged overall survival in patients who carried the EGFR mutation and received EGFR TKI targeted therapy”). Regarding claim 25, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 14 as outlined above. Further, Xiao teaches that the system of claim 24, wherein the operations further comprise generating a digitized IPD plot that conveys the IPD extracted from the digitized KM plot (See Xiao: Fig. 1, and [0107], “Although some deep-learning algorithms perform well in many preset computational challenges in the biomedical space, the performances of such algorithms often significantly deteriorate when applied to digital pathology images, due to the diversity of pathological images. For example, the differences between scanner type, manufacturer and digitization process can restrain the quality of digital pathology slides in terms of image sharpness, resolution, noise level, amplification magnitude, and/or the like. In the context of H&E-stained histopathology images analysis described herein, the color variation provides a unique challenge, which can be impacted by stain concentration, time elapsed, environmental temperatures upon staining, and/or other factors. Without properly accounting for these image quality issues and staining variations, classification, segmentation, and characterization may have decreased accuracy. In other words, inadequate image quality, low amplification magnitude, and staining variation may decrease an accuracy of tumor region segmentation, nuclei detection, and classification. Thus, the system 100 may perform image restoration and quality enhancement using the image restoration system 122 to restore blurred regions, enhance low resolution/magnification into high resolution, normalize staining colors to reduce staining variation, and/or the like”). Regarding claim 26, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 14 as outlined above. Further, Xiao teaches that the system of claim 14, wherein each KM curve of the one or more KM curves of the graphical KM plot depicts survival probability over time for a respective group of subjects (See Xiao: Fig. 14 and [0100], “The system 100 provides GCN-based pathology image analysis for histological classification in a variety of contexts. As described herein, in one example, the system 100 predicts responsiveness to EGFR TKI targeted therapy. Turning to FIG. 14, in one example, to predict a slide-level benefitting score, all image patches from the same pathology slide are grouped together to construct one disconnected graph. The GCN system 120 is trained to predict a benefitting score for each input graph and applied to patients with EGFR mutation in the testing set 118. As shown in a plot 1400 of survival proportion and time after metastasis, within the predicted benefitting group, patients who did not receive EGFR TKI targeted therapy showed significantly worse survival than patients who received EGFR TKI targeted therapy. As shown in the plot 1400, p=0.0002; without targeted therapy versus with targeted therapy, with a Hazard Ratio [HR]=6.81, and a 95% Confidence Interval [CI] 2.14-21.73. In contrast, within the predicted non-benefitting group, there was no significant survival difference between targeted therapy treated and non-treated patient groups. As shown in the plot 1400, p=0.10; without targeted therapy versus with targeted therapy, HR=2.22, and 95% CI 0.85-5.83. As further illustrated in Table 2, after adjusting for potential clinical confounders, including age, gender, smoking status, surgery, and stage at diagnosis, a high benefitting score calculated by the GCN system 120 is predictive for prolonged overall survival in patients who carried the EGFR mutation and received EGFR TKI targeted therapy”). Claims 8-9 and 21-22 are rejected under 35 U.S.C. 103 as being unpatentable over Lahmann. etc. (US 20170351708 A1) in view of Witte, etc. (US 20180253873 A1), further in view of Bekas, etc. (US 20190130614 A1) and Xiao, etc. (US 20230177682 A1), Xiao. etc. (US 20230177682 A1) and Siddiqi. etc. (US 20190272333 A1). Regarding claim 8, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 7 as outlined above. However, Lahmann, modified by Witte, Bekas and Xiao, fails to explicitly disclose that the method of claim 7, wherein processing each respective group of clustered pixels that represents the corresponding KM curve of the one or more KM curves of the graphical KM plot comprises: determining an average Euclidean distance between the pixels within the respective group of clustered pixels; identifying a top-N respective groups of clustered pixels having the lowest average Euclidean distances to represent each of the one or more KM curves; and processing the top-N respective groups of clustered pixels to generate the respective digitized representation of each corresponding KM curve of the one or more KM curves. However, Siddiqi teaches that the method of claim 7, wherein processing each respective group of clustered pixels that represents the corresponding KM curve of the one or more KM curves of the graphical KM plot (See Siddiqi: Fig. 1, and [0026], “FIG. 1 is a flow diagram that shows high-level operations of the data clustering method in accordance with an exemplary aspect of the disclosure. The data clustering method consists of two main operations. The input for the data clustering method S101 consists of the following items: (i) Set of data-points (D); (ii) Number of clusters (K), (iii) Five parameters (α, β, δ, p.sub.m, B). The first two parameters (α and β) belong to the first part and the remaining three parameters belong to the second part of the heuristic. The first main operation S103 is a greedy algorithm whose aim is to find points from the data-set that can act as centroids of clusters. A criteria for the selection of centroids is to maximize the inter-cluster separation. The second main operation S105 is a heuristic method that includes features found in GA and SimE algorithms and performs clustering by optimization. In one embodiment, the objective function of optimization is the validity index based on Calinski Harabasz index (CHI) or Dunn index (DI). The validity index ensures a solution, in S107, that optimizes both separation as well as compactness of clusters. The objective function may be represented by f.sub.n and its possible values are f.sub.n∈{CHI, DI}. The values of both indices are maximized in the optimization”. Note that the compactness is used to characterize each clustered group) comprises: determining an average Euclidean distance between the pixels within the respective group of clustered pixels (See Siddiqi: Fig. 1, and [0029], “The parameters α and β are related to the stopping criterion of the algorithm, where α represents the maximum number of iterations and β represents the number of iterations without changes. Being a greedy algorithm, the first main operation determines reasonable cluster centers that maximize separation between clusters in a manner that is more time-efficient than conventional evolutionary algorithms. Iterations in the method will be stopped when the value of α is reached or when β iterations have occurred without changes. In S301 of FIG. 3 and line 1 of the algorithm, initially up-to K data-points are selected as centroids. In line 4, set D holds the data-points which are not currently acting as centroids. In S305 and line 6, the set D.sub.z holds the centroids of all clusters except the z+1.sup.th cluster (the cluster C.sub.z is the z+1.sup.th cluster because the indices of clusters starts from zero.) The set P.sub.z stores a copy of the centroid of the z+1.sup.th cluster. In S307 and line 7, f.sub.0 is the value of the objective function before any change has taken place in the current iteration. In S309 and line 8, a data-point is chosen as the new centroid of the z+1.sup.th cluster. As the equation shows the new data-point should be the one which has maximum distance from the centroids of the remaining clusters. In S311 and line 9, the values of the objective function before and after the change are compared and, in step S315, the new centroid will be discarded if in S313 it worsens the value of the objective function. S317 and line 13 contains a condition to terminate the loop if the last β iterations are unable to produce any change in the centroids. In S303 and line 3, the algorithm can execute for up-to α number of iterations. In S319 and line 17, the first part of the algorithm returns the centers of K clusters”; and claim 2, “The computer system of claim 1, wherein the greedy search determines centroids of clusters that maximize inter-cluster separation”. Note that the inter-cluster separation is mapped to the Euclidean distance); identifying a top-N respective groups of clustered pixels having the lowest average Euclidean distances to represent each of the one or more KM curves (See Siddiqi: Fig. 1, and [0026], “FIG. 1 is a flow diagram that shows high-level operations of the data clustering method in accordance with an exemplary aspect of the disclosure. The data clustering method consists of two main operations. The input for the data clustering method S101 consists of the following items: (i) Set of data-points (D); (ii) Number of clusters (K), (iii) Five parameters (α, β, δ, p.sub.m, B). The first two parameters (α and β) belong to the first part and the remaining three parameters belong to the second part of the heuristic. The first main operation S103 is a greedy algorithm whose aim is to find points from the data-set that can act as centroids of clusters. A criteria for the selection of centroids is to maximize the inter-cluster separation. The second main operation S105 is a heuristic method that includes features found in GA and SimE algorithms and performs clustering by optimization. In one embodiment, the objective function of optimization is the validity index based on Calinski Harabasz index (CHI) or Dunn index (DI). The validity index ensures a solution, in S107, that optimizes both separation as well as compactness of clusters. The objective function may be represented by f.sub.n and its possible values are f.sub.n∈{CHI, DI}. The values of both indices are maximized in the optimization”. Note that the compactness is used to characterize each clustered group); and processing the top-N respective groups of clustered pixels to generate the respective digitized representation of each corresponding KM curve of the one or more KM curves (See Siddiqi: Fig. 1, and [0098], “If there is more than one data series comprised by the quantitative data each data series with a different type of markers, a clustering algorithm may be applied to map the markers to data series. A clustering algorithm groups a set of objects in such a way that objects in the same group, i.e. cluster, are more similar to each other regarding one or more predefined criteria, like e.g. their shape, than to those objects in other groups. The number of data series may be known, e.g. either from a legend or from a maximum number of markers detected per tick. Clustering techniques may be used on the marker shapes in order to identify which data series each marker belongs to”; Figs 4-5, and [0076], “In the second part of examples, the objective function can also be set to maximize cluster validity index DI. The results are presented in the same format as presented for CHI. Tables 7 and 8 present the solution quality (DI) and number of evaluations of the method and other algorithms. Table 7 shows solution quality results when objective function is to maximize DI. Table 8 shows a number of evaluations when objective function is to maximize DI. Tables 9 and 10 show the results of analysis using t-tests. Table 9 shows a comparison of the DI results of the heuristic with others using t-tests. The results in Table 9 convey the following information about the solution quality of the heuristic method: (i) It has better solution quality (DI) than Gen-SA in seven problems; (ii) It has a solution quality (DI) equal to Gen-SA in four problems; (iii) It is better than DE in solution quality (DI) in ten problems; (iv) It is equal to DE in three problems; (v) It is better than GA in ten problems; and (vi) It is equal to GA in two problems”; and claim 3 “The computer system of claim 2, wherein the greedy search selects a candidate centroid which has a maximum distance from determined centroids of the clusters and adds the candidate centroid to the determined centroid if the value of the validity index is improved”. Note that the candidate is searched and determined based on the validity index, and this is mapped to the top-N respective groups). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Gibson to have the method of claim 7, wherein processing each respective group of clustered pixels that represents the corresponding KM curve of the one or more KM curves of the graphical KM plot comprises: determining an average Euclidean distance between the pixels within the respective group of clustered pixels; identifying a top-N respective groups of clustered pixels having the lowest average Euclidean distances to represent each of the one or more KM curves; and processing the top-N respective groups of clustered pixels to generate the respective digitized representation of each corresponding KM curve of the one or more KM curves as taught by Siddiqi in order to avoid getting trapped in local optima and determines globally optimal solution (See Siddiqi: Fig. 1, and [0024], “The iterations continue until the stopping criterion (maximum runtime or maximum iterations) is reached. The heuristic method avoids getting trapped in local optima and determines globally optimal solution”). Lahmann teaches a method and system that may generate an alternative visualization of a data set based on a specification of a selected first visualization of the data set and parameters related to the data set; while Siddiqi teaches a system and method that may perform a greedy algorithm which selects the data-points that can act as centroids of clusters and the criterion to maximize the separation between clusters. Therefore, it is obvious for one of ordinary skill in the art to modify Lahmann by Siddiqi to perform a greedy algorithm which selects the data-points that can act as centroids of clusters and the criterion to maximize the separation between clusters. The motivation to modify Lahmann by Siddiqi is “Simple substitution of one known element for another to obtain predictable results”. Regarding claim 9, Lahmann, Witte, Bekas, Xiao and Siddiqi teach all the features with respect to claim 8 as outlined above. Further, Siddiqi teaches that the method of claim 8, wherein N is equal to a number of the one or more KM curves of the graphical KM plot (See Siddiqi: Fig. 1, and [0026], “FIG. 1 is a flow diagram that shows high-level operations of the data clustering method in accordance with an exemplary aspect of the disclosure. The data clustering method consists of two main operations. The input for the data clustering method S101 consists of the following items: (i) Set of data-points (D); (ii) Number of clusters (K), (iii) Five parameters (α, β, δ, p.sub.m, B). The first two parameters (α and β) belong to the first part and the remaining three parameters belong to the second part of the heuristic. The first main operation S103 is a greedy algorithm whose aim is to find points from the data-set that can act as centroids of clusters. A criteria for the selection of centroids is to maximize the inter-cluster separation. The second main operation S105 is a heuristic method that includes features found in GA and SimE algorithms and performs clustering by optimization. In one embodiment, the objective function of optimization is the validity index based on Calinski Harabasz index (CHI) or Dunn index (DI). The validity index ensures a solution, in S107, that optimizes both separation as well as compactness of clusters. The objective function may be represented by f.sub.n and its possible values are f.sub.n∈{CHI, DI}. The values of both indices are maximized in the optimization.”; and [0006], “In optimization perspective, clustering problem is considered as an NP-hard grouping problem”. Note that in the optimization process to maximize the validity index in clustering the colors of the image having a know number of distinct colored object (here is number of KM curves), setting L the number of clusters) equal to the number of those colored objects is the ordinary and expected choices). Regarding claim 21, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 20 as outlined above. Further, Lahmann teaches that the system of claim 20, wherein processing each respective group of clustered pixels that represents the corresponding KM curve of the one or more KM curves of the graphical KM plot (See Siddiqi: Fig. 1, and [0026], “FIG. 1 is a flow diagram that shows high-level operations of the data clustering method in accordance with an exemplary aspect of the disclosure. The data clustering method consists of two main operations. The input for the data clustering method S101 consists of the following items: (i) Set of data-points (D); (ii) Number of clusters (K), (iii) Five parameters (α, β, δ, p.sub.m, B). The first two parameters (α and β) belong to the first part and the remaining three parameters belong to the second part of the heuristic. The first main operation S103 is a greedy algorithm whose aim is to find points from the data-set that can act as centroids of clusters. A criteria for the selection of centroids is to maximize the inter-cluster separation. The second main operation S105 is a heuristic method that includes features found in GA and SimE algorithms and performs clustering by optimization. In one embodiment, the objective function of optimization is the validity index based on Calinski Harabasz index (CHI) or Dunn index (DI). The validity index ensures a solution, in S107, that optimizes both separation as well as compactness of clusters. The objective function may be represented by f.sub.n and its possible values are f.sub.n∈{CHI, DI}. The values of both indices are maximized in the optimization”. Note that the compactness is used to characterize each clustered group) comprises: determining an average Euclidean distance between the pixels within the respective group of clustered pixels (See Siddiqi: Fig. 1, and [0029], “The parameters α and β are related to the stopping criterion of the algorithm, where α represents the maximum number of iterations and β represents the number of iterations without changes. Being a greedy algorithm, the first main operation determines reasonable cluster centers that maximize separation between clusters in a manner that is more time-efficient than conventional evolutionary algorithms. Iterations in the method will be stopped when the value of α is reached or when β iterations have occurred without changes. In S301 of FIG. 3 and line 1 of the algorithm, initially up-to K data-points are selected as centroids. In line 4, set D holds the data-points which are not currently acting as centroids. In S305 and line 6, the set D.sub.z holds the centroids of all clusters except the z+1.sup.th cluster (the cluster C.sub.z is the z+1.sup.th cluster because the indices of clusters starts from zero.) The set P.sub.z stores a copy of the centroid of the z+1.sup.th cluster. In S307 and line 7, f.sub.0 is the value of the objective function before any change has taken place in the current iteration. In S309 and line 8, a data-point is chosen as the new centroid of the z+1.sup.th cluster. As the equation shows the new data-point should be the one which has maximum distance from the centroids of the remaining clusters. In S311 and line 9, the values of the objective function before and after the change are compared and, in step S315, the new centroid will be discarded if in S313 it worsens the value of the objective function. S317 and line 13 contains a condition to terminate the loop if the last β iterations are unable to produce any change in the centroids. In S303 and line 3, the algorithm can execute for up-to α number of iterations. In S319 and line 17, the first part of the algorithm returns the centers of K clusters”; and claim 2, “The computer system of claim 1, wherein the greedy search determines centroids of clusters that maximize inter-cluster separation”. Note that the inter-cluster separation is mapped to the Euclidean distance); identifying a top-N respective groups of clustered pixels having the lowest average Euclidean distances to represent each of the one or more KM curves (See Siddiqi: Fig. 1, and [0026], “FIG. 1 is a flow diagram that shows high-level operations of the data clustering method in accordance with an exemplary aspect of the disclosure. The data clustering method consists of two main operations. The input for the data clustering method S101 consists of the following items: (i) Set of data-points (D); (ii) Number of clusters (K), (iii) Five parameters (α, β, δ, p.sub.m, B). The first two parameters (α and β) belong to the first part and the remaining three parameters belong to the second part of the heuristic. The first main operation S103 is a greedy algorithm whose aim is to find points from the data-set that can act as centroids of clusters. A criteria for the selection of centroids is to maximize the inter-cluster separation. The second main operation S105 is a heuristic method that includes features found in GA and SimE algorithms and performs clustering by optimization. In one embodiment, the objective function of optimization is the validity index based on Calinski Harabasz index (CHI) or Dunn index (DI). The validity index ensures a solution, in S107, that optimizes both separation as well as compactness of clusters. The objective function may be represented by f.sub.n and its possible values are f.sub.n∈{CHI, DI}. The values of both indices are maximized in the optimization”. Note that the compactness is used to characterize each clustered group); and processing the top-N respective groups of clustered pixels to generate the respective digitized representation of each corresponding KM curve of the one or more KM curves . (See Siddiqi: Fig. 1, and [0098], “If there is more than one data series comprised by the quantitative data each data series with a different type of markers, a clustering algorithm may be applied to map the markers to data series. A clustering algorithm groups a set of objects in such a way that objects in the same group, i.e. cluster, are more similar to each other regarding one or more predefined criteria, like e.g. their shape, than to those objects in other groups. The number of data series may be known, e.g. either from a legend or from a maximum number of markers detected per tick. Clustering techniques may be used on the marker shapes in order to identify which data series each marker belongs to”; Figs 4-5, and [0076], “In the second part of examples, the objective function can also be set to maximize cluster validity index DI. The results are presented in the same format as presented for CHI. Tables 7 and 8 present the solution quality (DI) and number of evaluations of the method and other algorithms. Table 7 shows solution quality results when objective function is to maximize DI. Table 8 shows a number of evaluations when objective function is to maximize DI. Tables 9 and 10 show the results of analysis using t-tests. Table 9 shows a comparison of the DI results of the heuristic with others using t-tests. The results in Table 9 convey the following information about the solution quality of the heuristic method: (i) It has better solution quality (DI) than Gen-SA in seven problems; (ii) It has a solution quality (DI) equal to Gen-SA in four problems; (iii) It is better than DE in solution quality (DI) in ten problems; (iv) It is equal to DE in three problems; (v) It is better than GA in ten problems; and (vi) It is equal to GA in two problems”; and claim 3 “The computer system of claim 2, wherein the greedy search selects a candidate centroid which has a maximum distance from determined centroids of the clusters and adds the candidate centroid to the determined centroid if the value of the validity index is improved”. Note that the candidate is searched and determined based on the validity index, and this is mapped to the top-N respective groups) Regarding claim 22, Lahmann, Witte, Bekas, Xiao and Siddiqi teach all the features with respect to claim 21 as outlined above. Further, Siddiqi teaches that the system of claim 21, wherein N is equal to a number of the one or more KM curves of the graphical KM plot (See Siddiqi: Fig. 1, and [0026], “FIG. 1 is a flow diagram that shows high-level operations of the data clustering method in accordance with an exemplary aspect of the disclosure. The data clustering method consists of two main operations. The input for the data clustering method S101 consists of the following items: (i) Set of data-points (D); (ii) Number of clusters (K), (iii) Five parameters (α, β, δ, p.sub.m, B). The first two parameters (α and β) belong to the first part and the remaining three parameters belong to the second part of the heuristic. The first main operation S103 is a greedy algorithm whose aim is to find points from the data-set that can act as centroids of clusters. A criteria for the selection of centroids is to maximize the inter-cluster separation. The second main operation S105 is a heuristic method that includes features found in GA and SimE algorithms and performs clustering by optimization. In one embodiment, the objective function of optimization is the validity index based on Calinski Harabasz index (CHI) or Dunn index (DI). The validity index ensures a solution, in S107, that optimizes both separation as well as compactness of clusters. The objective function may be represented by f.sub.n and its possible values are f.sub.n∈{CHI, DI}. The values of both indices are maximized in the optimization.”; and [0006], “In optimization perspective, clustering problem is considered as an NP-hard grouping problem”. Note that in the optimization process to maximize the validity index in clustering the colors of the image having a know number of distinct colored object (here is number of KM curves), setting L the number of clusters) equal to the number of those colored objects is the ordinary and expected choices). Claims 10 and 23 are rejected under 35 U.S.C. 103 as being unpatentable over Lahmann. etc. (US 20170351708 A1) in view of Witte, etc. (US 20180253873 A1), further in view of Bekas, etc. (US 20190130614 A1) and Xiao, etc. (US 20230177682 A1), Xiao. etc. (US 20230177682 A1), Siddiqi. etc. (US 20190272333 A1) and Jones (US 20180329865 A1). Regarding claim 10, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 1 as outlined above. However, Lahmann, modified by Witte, Bekas, Xiao and Siddiqi, fails to explicitly disclose that the method of claim 1, wherein generating the digitized KM plot comprises, for each corresponding pixel point in the respective digitized representation generated for the corresponding KM curve: applying a regression model over a window encompassing a predetermined number pixels before and a the predetermined number of pixels after the corresponding pixel point to estimate a reference y value for the corresponding pixel point; calculating a mean and standard deviation of discrepancies between the reference y value and the counterpart digitized y value extracted from the respective digitized representation; and removing the corresponding pixel point from the respective digitized representation when the standard deviation from the mean exceeds a threshold number of standard deviations. However, Jones teaches that the method of claim 1, wherein generating the digitized KM plot comprises, for each corresponding pixel point in the respective digitized representation generated for the corresponding KM curve (See Jones: Figs. 1-3, and [0010], “Another embodiment includes a computer implemented method for assessing the viability of a data set as used in developing a model comprising the steps of: providing a target data set comprising a plurality of data values; generating a random target data set based on the target dataset; selecting a set of bias criteria values; generating, by a processor, an outlier bias reduced target data set based on the data set and each of the selected bias criteria values; generating, by the processor, an outlier bias reduced random data set based on the random data set and each of the selected bias criteria values; calculating a set of error values for the outlier bias reduced data set and the outlier bias reduced random data set; calculating a set of correlation coefficients for the outlier bias reduced data set and the outlier bias reduced random data set; generating bias criteria curves for the data set and the random data set based on the selected bias criteria values and the corresponding error value and correlation coefficient; and comparing the bias criteria curve for the data set to the bias criteria curve for the random data set. The outlier bias reduced target data set and the outlier bias reduced random target data set are generated using the Dynamic Outlier Bias Removal methodology. The random target data set can comprise of randomized data values developed from values within the range of the plurality of data values. Also, the set of error values can comprise a set of standard errors, and wherein the set of correlation coefficients comprises a set of coefficient of determination values. Another embodiment can further comprise the step of generating automated advice regarding the viability of the target data set to support the developed model, and vice versa, based on comparing the bias criteria curve for the target data set to the bias criteria curve for the random target data set. Advice can be generated based on parameters selected by analysts, such as a correlation coefficient threshold and/or an error threshold. Yet another embodiment further comprises the steps of: providing an actual data set comprising a plurality of actual data values corresponding to the model predicted values; generating a random actual data set based on the actual data set; generating, by a processor, an outlier bias reduced actual data set based on the actual data set and each of the selected bias criteria values; generating, by the processor, an outlier bias reduced random actual data set based on the random actual data set and each of the selected bias criteria values; generating, for each selected bias criteria, a random data plot based on the outlier bias reduced random target data set and the outlier bias reduced random actual data; generating, for each selected bias criteria, a realistic data plot based on the outlier bias reduced target data set and the outlier bias reduced actual target data set; and comparing the random data plot with the realistic data plot corresponding to each of the selected bias criteria.”): applying a regression model over a window encompassing a predetermined number pixels before and a the predetermined number of pixels after the corresponding pixel point to estimate a reference y value for the corresponding pixel point (See Jones: Figs. 1-3, and [0073], “] In one embodiment, Dynamic Outlier Bias Reduction is applied to a procedure that uses the data and a prescribed overall error criteria to determine statistical outliers that are removed from the model coefficient calculations. This is a data-driven process that identifies outliers using a data produced global error criteria using for example, the percentile function. The use of Dynamic Outlier Bias Reduction is not limited to the reduction of bias in model predicted values, and its use in this embodiment is illustrative and exemplary only. Dynamic Outlier Bias Reduction may also be used, for example, to remove outliers from any statistical data set, including use in calculation of, but not limited to, arithmetic averages, linear regressions, and trend lines. The outlier facilities are still ranked from the calculation results, but the outliers are not used in the filtered data set applied to compute model coefficients or statistical results”); calculating a mean and standard deviation of discrepancies between the reference y value and the counterpart digitized y value extracted from the respective digitized representation (See Jones: Figs. 1-3, and [0074], “A standard procedure, commonly used to remove outliers, is to compute the standard deviation (σ) of the data set and simply define all data outside a 2σ interval of the mean, for example, as outliers. This procedure has statistical assumptions that, in general, cannot be tested in practice. The Dynamic Outlier Bias Reduction method description applied in an embodiment of this invention, is outlined in FIG. 1, uses both a relative error and absolute error. For example: for a facility, ‘m’:”); and removing the corresponding pixel point from the respective digitized representation when the standard deviation from the mean exceeds a threshold number of standard deviations (See Jones: Figs. 1-3, and [0074], “A standard procedure, commonly used to remove outliers, is to compute the standard deviation (σ) of the data set and simply define all data outside a 2σ interval of the mean, for example, as outliers. This procedure has statistical assumptions that, in general, cannot be tested in practice. The Dynamic Outlier Bias Reduction method description applied in an embodiment of this invention, is outlined in FIG. 1, uses both a relative error and absolute error. For example: for a facility, ‘m’:” Note thet 2 standard deviation is mapped to the threshold). Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention was effectively filed to modify Gibson to have the method of claim 1, wherein generating the digitized KM plot comprises, for each corresponding pixel point in the respective digitized representation generated for the corresponding KM curve: applying a regression model over a window encompassing a predetermined number pixels before and a the predetermined number of pixels after the corresponding pixel point to estimate a reference y value for the corresponding pixel point; calculating a mean and standard deviation of discrepancies between the reference y value and the counterpart digitized y value extracted from the respective digitized representation; and removing the corresponding pixel point from the respective digitized representation when the standard deviation from the mean exceeds a threshold number of standard deviations as taught by Jones in order to enable removing outlier data bias using a dynamic statistical process useful for data quality operations, data validation, statistic calculations or mathematical model development. (See Jones: Fig. 1, and [0002], “Removing outlier data in standards or data-driven model development is an important part of the pre-analysis work to ensure a representative and fair analysis is developed from the underlying data. For example, developing equitable benchmarking of greenhouse gas standards for carbon dioxide (CO.sub.2), ozone (O.sub.3), water vapor (H.sub.2O), hydrofluorocarbons (HFCs), perfluorocarbons (PFCs), chlorofluorocarbons (CFCs), sulfur hexafluoride (SF.sub.6), methane (CH.sub.4), nitrous oxide (N.sub.2O), carbon monoxide (CO), nitrogen oxides (NOx), and non-methane volatile organic compounds (NMVOCs) emissions requires that collected industrial data used in the standards development exhibit certain properties”). Lahmann teaches a method and system that may generate an alternative visualization of a data set based on a specification of a selected first visualization of the data set and parameters related to the data set; while Jones teaches a system and method that may filter data to reduce functional, and trend line outlier bias by remove outliers from the data set through an objective statistical method. Therefore, it is obvious for one of ordinary skill in the art to modify Lahmann by Jones to filter data to reduce functional, and trend line outlier bias by remove outliers from the data set through an objective statistical method. The motivation to modify Lahmann by Jones is “Simple substitution of one known element for another to obtain predictable results”. Regarding claim 23, Lahmann, Witte, Bekas and Xiao teach all the features with respect to claim 14 as outlined above. Further, Jones teaches that the system of claim 14, wherein generating the digitized KM plot comprises, for each corresponding pixel point in the respective digitized representation generated for the corresponding KM curve(See Jones: Figs. 1-3, and [0010], “Another embodiment includes a computer implemented method for assessing the viability of a data set as used in developing a model comprising the steps of: providing a target data set comprising a plurality of data values; generating a random target data set based on the target dataset; selecting a set of bias criteria values; generating, by a processor, an outlier bias reduced target data set based on the data set and each of the selected bias criteria values; generating, by the processor, an outlier bias reduced random data set based on the random data set and each of the selected bias criteria values; calculating a set of error values for the outlier bias reduced data set and the outlier bias reduced random data set; calculating a set of correlation coefficients for the outlier bias reduced data set and the outlier bias reduced random data set; generating bias criteria curves for the data set and the random data set based on the selected bias criteria values and the corresponding error value and correlation coefficient; and comparing the bias criteria curve for the data set to the bias criteria curve for the random data set. The outlier bias reduced target data set and the outlier bias reduced random target data set are generated using the Dynamic Outlier Bias Removal methodology. The random target data set can comprise of randomized data values developed from values within the range of the plurality of data values. Also, the set of error values can comprise a set of standard errors, and wherein the set of correlation coefficients comprises a set of coefficient of determination values. Another embodiment can further comprise the step of generating automated advice regarding the viability of the target data set to support the developed model, and vice versa, based on comparing the bias criteria curve for the target data set to the bias criteria curve for the random target data set. Advice can be generated based on parameters selected by analysts, such as a correlation coefficient threshold and/or an error threshold. Yet another embodiment further comprises the steps of: providing an actual data set comprising a plurality of actual data values corresponding to the model predicted values; generating a random actual data set based on the actual data set; generating, by a processor, an outlier bias reduced actual data set based on the actual data set and each of the selected bias criteria values; generating, by the processor, an outlier bias reduced random actual data set based on the random actual data set and each of the selected bias criteria values; generating, for each selected bias criteria, a random data plot based on the outlier bias reduced random target data set and the outlier bias reduced random actual data; generating, for each selected bias criteria, a realistic data plot based on the outlier bias reduced target data set and the outlier bias reduced actual target data set; and comparing the random data plot with the realistic data plot corresponding to each of the selected bias criteria.”): applying a regression model over a window encompassing a predetermined number pixels before and a the predetermined number of pixels after the corresponding pixel point to estimate a reference y value for the corresponding pixel point (See Jones: Figs. 1-3, and [0073], “] In one embodiment, Dynamic Outlier Bias Reduction is applied to a procedure that uses the data and a prescribed overall error criteria to determine statistical outliers that are removed from the model coefficient calculations. This is a data-driven process that identifies outliers using a data produced global error criteria using for example, the percentile function. The use of Dynamic Outlier Bias Reduction is not limited to the reduction of bias in model predicted values, and its use in this embodiment is illustrative and exemplary only. Dynamic Outlier Bias Reduction may also be used, for example, to remove outliers from any statistical data set, including use in calculation of, but not limited to, arithmetic averages, linear regressions, and trend lines. The outlier facilities are still ranked from the calculation results, but the outliers are not used in the filtered data set applied to compute model coefficients or statistical results”); calculating a mean and standard deviation of discrepancies between the reference y value and the counterpart digitized y value extracted from the respective digitized representation (See Jones: Figs. 1-3, and [0074], “A standard procedure, commonly used to remove outliers, is to compute the standard deviation (σ) of the data set and simply define all data outside a 2σ interval of the mean, for example, as outliers. This procedure has statistical assumptions that, in general, cannot be tested in practice. The Dynamic Outlier Bias Reduction method description applied in an embodiment of this invention, is outlined in FIG. 1, uses both a relative error and absolute error. For example: for a facility, ‘m’:”); and removing the corresponding pixel point from the respective digitized representation when the standard deviation from the mean exceeds a threshold number of standard deviations (See Jones: Figs. 1-3, and [0074], “A standard procedure, commonly used to remove outliers, is to compute the standard deviation (σ) of the data set and simply define all data outside a 2σ interval of the mean, for example, as outliers. This procedure has statistical assumptions that, in general, cannot be tested in practice. The Dynamic Outlier Bias Reduction method description applied in an embodiment of this invention, is outlined in FIG. 1, uses both a relative error and absolute error. For example: for a facility, ‘m’:” Note thet 2 standard deviation is mapped to the threshold). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to GORDON G LIU whose telephone number is (571)270-0382. The examiner can normally be reached Monday - Friday 8:00-5:00. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona E Faulk can be reached at 571-272-7515. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /GORDON G LIU/ Primary Examiner, Art Unit 2618
Read full office action

Prosecution Timeline

Mar 28, 2025
Application Filed
Sep 23, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743838
PIXEL GENERATION TECHNIQUE
2y 8m to grant Granted Sep 22, 2026
Patent 12743821
RECORDING MEDIUM AND INFORMATION PROCESSING DEVICE
2y 3m to grant Granted Sep 22, 2026
Patent 12744018
IMAGE OUTPUT CONTROL DEVICE AND METHOD
2y 1m to grant Granted Sep 22, 2026
Patent 12725544
SCREEN DISPLAY DRIVING METHOD AND APPARATUS, SCREEN INFORMATION CONFIGURATION METHOD AND APPARATUS, MEDIUM AND DEVICE
2y 1m to grant Granted Sep 01, 2026
Patent 12725367
PICTURE DISPLAY METHOD, SYSTEM, AND APPARATUS, DEVICE, AND STORAGE MEDIUM
2y 2m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
83%
Grant Probability
98%
With Interview (+14.8%)
2y 2m (~7m remaining)
Median Time to Grant
Low
PTA Risk
Based on 701 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month