DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-3 and 13-20 are rejected under 35 U.S.C. 103 as being unpatentable over SETLUR (No. US-11822584-B2 “Setlur”) in view of KEAHEY (No. US-20210157819-A1 “Keahey”) and in further view of JAIN (No. US-20230267175-A1 “Jain”).
Regarding claim 1, Setlur teaches “A method for visualizing large datasets, performed by a computing device executing a browser application, the method comprising:” (a method executes at a computing device; Col 2, Line 20-21); (a method for generating data visualizations; Col 4, Line 17-18); (executed by a web browser on a user's computing device; Col 9, Line 34-35);
“Obtaining a dataset for rendering a data visualization” (from the data source – Col 2, lines 36-40); “the data set including a plurality of data point” (for example: the price points for the wine varieties- Col 11, lines 25-39- figure 3A- “Once a data source is selected, the data visualization application 230 generates a visual specification that specifies the selected data source, a plurality of visual variables, and a plurality of data fields from the data source. Each of the visual variables is associated with a respective one or more of the data fields and each of the data fields is identified as either a dimension or a measure. The visual variables include information that encode how the data visualization will look (e.g., data visualization type, what data points will be displayed or represented as visual marks, the color scheme of visual marks, or emphasizing certain visual marks). A data visualization is generated and displayed based of the visual specification”);
“Selecting, from the plurality of data points, a first subset” (for example: expensive wine varieties) “of data paints (col. 11, lines 40-64: “For example, a user may ask, “expensive varieties.” In this example, “expensive” is a superlative adjective indicating that the user may want to see the wine varieties (for the running example) that have the highest cost. In some implementations, in response to the natural language command, the data visualization application 230 identifies a first keyword (e.g., “varieties”) in the natural language command and one or more second keywords (e.g., “expensive”) in the natural language command that are adjectives that modify the first keyword. The data visualization then generates a visual specification or modifies an existing visual specification so that the second keyword corresponds to one or more first data fields of the plurality of data fields (e.g., select the “Price” data field). One or more visual variables are associated with the one or more first data fields according to the one or more second keywords (e.g., visual variables associated with filtering or emphasizing/deemphasizing is associated with a data field corresponding to a number of patients by age so that the age bins that have the greatest number of records are emphasized/highlighted or shown). The data visualization application 230 then generates a data visualization in accordance with (e.g., based on) the visual specification and displays the data visualization in the data visualization region 112.;
While Setlar does not teach “selecting, from the plurality of data points, a first subset of data paints according to a statistical data distribution of the data set” and “recursively applying a first algorithm to the first subset of data points to obtain a final subset of data points, wherein each of first subset of data points and the final subset of data points has a fewer number of data points than the plurality of data points”.
Setlur teaches “rendering a data visualization using the browser application, the data visualization having a plurality of data marks corresponding to the final subset of data points; and” (generates and displays a data visualization, including a plurality of visual marks representing data retrieved from the data source; Col 2, Line 40-42); (the data visualization application 230 executes within the web browser 226 or another application using web pages provided by a web server; Col 7, Line 62-65);
“displaying, on the browser application, the data visualization including the plurality of data marks.” (generates and displays a data visualization, including a plurality of visual marks representing data retrieved from the data source; Col 2, Line 40-42); (the data visualization application 230 executes within the web browser 226 or another application using web pages provided by a web server; Col 7, Line 62-65);
As stated above, Setlur does not teach “selecting, from the plurality of data points, a first subset of data points according to a statistical data distribution of the dataset”.
Keahey teaches “selecting, from the plurality of data points, a first subset of data points according to a statistical data distribution of the dataset;” (determining the distribution ... by performing statistical analysis .... determining a collection of the data visualizations based on the statistical analysis performed on each of the vectors, the collection of the data visualizations comprising at least a subset; Para 0005);
It would be obvious to a person skilled in the art to perform statistical analysis to get a subset of data points. When you obtain a dataset that contains a plurality of data points, a standard practice in data visualization is statistical data distribution of the dataset. Performing statistical analysis would get you a subset used for further operations of data visualization.
Thus, it would have been obvious to one of ordinary skill in the art at the effective filing date of the claimed invention to modify Setlur by selecting, from the plurality of data points, a first subset of data points according to a statistical data distribution of the dataset as taught by Keahey.
Setlur and Keahey fail to specifically disclose “recursively applying a first algorithm to the first subset of data points to obtain a final subset of data points, wherein each of first subset of data points and the final subset of data points has a fewer number of data points than the plurality of data points”.
Jain teaches “recursively applying a first algorithm to the first subset of data points to obtain a final subset of data points, wherein each of first subset of data points and the final subset of data points has a fewer number of data points than the plurality of data points;” (in an iterative manner, with the input dataset for a next iteration comprising an augmented set of labeled examples from a current iteration and selected unlabeled examples, until a final set of labeled examples is created; Abstract);
Setlur discloses the method and generation or rendering of the visualization on an application browser which teaches the claimed subject matter. Keahey teaches getting a subset based on the performing of statistical analysis which teaches the claimed subject matter of a first subset of data points according to a statistical data distribution. Jain discloses a dataset with data points as well as iterative manner to result in a final subset of data points. This teaches the claimed subject matter of obtaining a dataset and recursively applying an algorithm to obtain a final subset of data points. In combination, Setlur, Keahey and Jain teaches the visualizing large datasets.
Setlur, Keahey and Jain are analogous art as they are related to data visualization and datasets.
The motivation for the above is to have accurate visualization of datasets for easier user application and usage.
It would be obvious to a person skilled in the art to recursively apply the algorithm to the subset of data points to get a final subset of data points. A standard practice in data visualization is iteratively applying an algorithm to produce the final data points used to generate the final distribution of the data visualization.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur in view of Keahey by recursively applying a first algorithm to the first subset of data points to obtain a final subset of data points, wherein each of first subset of data points and the final subset of data points has a fewer number of data points than the plurality of data points as taught by Jain.
Regarding claim 2, Jain further teaches “The method of claim 1, wherein recursively applying the first algorithm to the first subset of data points to obtain the final subset of data points includes:
applying the first algorithm to the first subset of data points to obtain a second subset of data points;” (multiple iterations of mapping an input dataset that contains a set of labeled training data 52 (potentially as augmented by prior iterations) and a subset of unlabeled training data 54 to a reduced-dimension data space, using the reduced-dimension data space to identify data points; Para 00057);
“dividing the second subset of data points into multiple data segments, each of the data segments including a respective third subset of data points; and” (multiple iterations of mapping an input dataset that contains a set of labeled training data 52 (potentially as augmented by prior iterations) and a subset of unlabeled training data 54 to a reduced-dimension data space, using the reduced-dimension data space to identify data points; Para 00057);
“reapplying the first algorithm to a least a portion of each data segment, of the multiple data segments, to obtain a respective fourth subset of data points from the respective third subset of data points.” (multiple iterations of mapping an input dataset that contains a set of labeled training data 52 (potentially as augmented by prior iterations) and a subset of unlabeled training data 54 to a reduced-dimension data space, using the reduced-dimension data space to identify data points; Para 00057); (can output a final set of labeled data, which includes labeled examples generated through the iterative process. The final set of labeled data can be input to ML model training component 118, such as an ML training algorithm; Para 0104);
Jain discloses repeated iterative application to subsets to get final data points. This teaches the claimed subject matter of applying algorithm to subset, obtaining another subset, and reapplying segments until final subset. It is obvious to a person skilled in the art that an iterative process will give a first subset of data, second, third, fourth, and so on until the condition is met, which is the final data points.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur in view of Keahey by applying the first algorithm to the first subset of data points to obtain a second subset of data points, dividing the second subset of data points into multiple data segments, each of the data segments including a respective third subset of data points and reapplying the first algorithm to a least a portion of each data segment, of the multiple data segments, to obtain a respective fourth subset of data points from the respective third subset of data points as taught by Jain.
Regarding claim 3, Jain further teaches “The method of claim 2, wherein reapplying the first algorithm to the a least a portion of each data segment to obtain the respective fourth subset of data points:
for each data segment:
determining a respective tolerance value for the data segment according to characteristics of the respective fourth subset of data points;” (The iterative process of labeling training data can continue until a stopping condition is met, such as performing a certain number of iterations, collecting a threshold amount of labeled training data 56; Para 0060);
“in accordance with a determination that the respective fourth subset of data points satisfy the respective tolerance value:” (The iterative process of labeling training data can continue until a stopping condition is met, such as performing a certain number of iterations, collecting a threshold amount of labeled training data 56 or satisfying another criterion; Para 0060);
“retaining the respective fourth subset of data points; and
including the respective fourth subset of data points in the final subset of final points; and” (multiple iterations of mapping an input dataset that contains a set of labeled training data 52 (potentially as augmented by prior iterations) and a subset of unlabeled training data 54 to a reduced-dimension data space, using the reduced-dimension data space to identify data points; Para 00057); (can output a final set of labeled data, which includes labeled examples generated through the iterative process. The final set of labeled data can be input to ML model training component 118, such as an ML training algorithm; Para 0104);
“in accordance with a determination that the respective fourth subset of data points do not satisfy the respective tolerance value:” (The iterative process of labeling training data can continue until a stopping condition is met, such as performing a certain number of iterations, collecting a threshold amount of labeled training data 56 or satisfying another criterion; Para 0060);
“dividing the data segment into one or more sub-segments; and
reapplying the first algorithm to each of the sub-segments.” (multiple iterations of mapping an input dataset that contains a set of labeled training data 52 (potentially as augmented by prior iterations) and a subset of unlabeled training data 54 to a reduced-dimension data space, using the reduced-dimension data space to identify data points; Para 00057);
Jain further discloses the stopping criteria and further iteration process. This teaches the claimed subject matter retaining a subset when it satisfies the tolerance/criterion and dividing and reapplying when it does not. It is obvious to a person skilled in the art that an iterative process will give a first subset of data, second, third, fourth, and so on until the condition is met, which is the final data points. In an iterative process, if the condition is not met, it will continue the iterative process by reapplying the algorithm.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur in view of Keahey by reapplying the first algorithm to the a least a portion of each data segment to obtain the respective fourth subset of data points: for each data segment: determining a respective tolerance value for the data segment according to characteristics of the respective fourth subset of data points, in accordance with a determination that the respective fourth subset of data points satisfy the respective tolerance value: retaining the respective fourth subset of data points; and including the respective fourth subset of data points in the final subset of final points; and in accordance with a determination that the respective fourth subset of data points do not satisfy the respective tolerance value dividing the data segment into one or more sub-segments; and reapplying the first algorithm to each of the sub-segments as taught by Jain.
Regarding claim 13, Jain further teaches “The method of claim 1, further comprising:
after obtaining the dataset and prior to selecting the first subset of data points:
performing feature extraction on the dataset to identify, from the plurality of data points, an initial subset of data points that retains a visual perception of the data visualization.” (may be represented by a respective high dimension feature vector; Para 0047); (the data points in area 504 ... are mapped to the reduced-dimension representation of the selected area for evaluation by the selection model. Thus, the selection model may identify target data points; Para 0117);
Jain discloses a high dimension feature vector that showcases the feature extraction. The features are processed through reduced dimension to produce the visualization space. Thus, the selection model, before it selects, first it identifies target examples from the reduced dimension representation. This teaches the initial subset of data points that retains a visual perception of the claimed subject matter.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur in view of Keahey by after obtaining the dataset and prior to selecting the first subset of data points: performing feature extraction on the dataset to identify, from the plurality of data points, an initial subset of data points that retains a visual perception of the data visualization as taught by Jain.
Regarding claim 14, Setlur teaches “The method of claim 1, wherein the selecting the first subset of data points is further based on a data mark encoding type of the data visualization.” (visual variables include information that encode how the data visualization will look (e.g., data visualization type, what data points will be displayed or represented as visual marks; Col 11, Line 33-36);
Setlur discloses displaying or representing based on visual encoding which can teach the selection of data points based on the encoding of the mark.
Regarding claim 15, Keahey further teaches “The method of claim 1, wherein selecting, from the plurality of data points, the first subset of data points according to the data distribution of the dataset includes:
applying a machine learning model to determine, from the statistical data distribution, the first subset of data points such that the first subset of data points preserves a visual perception of the data visualization.” (operations also can include determining the distribution of the data fields and the data visualization channels across the plurality of data visualizations by performing statistical analysis on each of the vectors. The operations also can include determining a collection of the data visualizations based on the statistical analysis performed on each of the vectors, the collection of the data visualizations comprising at least a subset of the plurality of data visualizations; Para 0005); (training system, a series of data may be collected that relates to the specified criteria and/or characteristics; Para 0047); (the visualization evaluation module can identify data visualization channels used in the first data visualization. Data visualization style, such as size, color, and type influencing how a particular data visualization channel may look, may be adjusted based on switches applied to particular data visualization channels; Para 0087);
Keahey discloses performing statistical analysis to determine the distribution of the dataset. The training system uses statistical distribution to determine a subset collection of data. As well as disclosing the visualization style that affect how the visualization looks. This teaches the claimed subject matter of from the statistical data distribution, the first subset of data points preserve a visual perception of the data visualization
The motivation for the above is to accurate visual perception of data visualization for accurate user viewing.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur by applying a machine learning model to determine, from the statistical data distribution, the first subset of data points such that the first subset of data points preserves a visual perception of the data visualization as taught by Keahey.
Regarding claim 16, Keahey further teaches “The method of claim 1, wherein selecting, from the plurality of data points, the first subset of data points according to the data distribution of the dataset includes:
applying a machine learning model to determine, from the statistical data distribution, a second subset of data points from the plurality of data points; and” (operations also can include determining the distribution of the data fields and the data visualization channels across the plurality of data visualizations by performing statistical analysis on each of the vectors. The operations also can include determining a collection of the data visualizations based on the statistical analysis performed on each of the vectors, the collection of the data visualizations comprising at least a subset of the plurality of data visualizations; Para 0005); (training system, a series of data may be collected that relates to the specified criteria and/or characteristics; Para 0047); (visualization evaluation module can identify data visualization channels used in the second data visualization; Para 0087);
“performing a filtering or grouping operation on each data point of the second subset of data points.” (the visualization evaluation module can consider various other data visualization properties, such as whether data aggregation/binning, filtering, or disaggregation, have been applied; Para 0082); (visualization properties may also influence how data is presented, such as whether data has been aggregated or binned, filtered, or disaggregated; Para 0087);
Keahey discloses performing statistical analysis to determine the distribution of the dataset. The training system uses statistical distribution to determine a subset collection of data, thus the evaluation module can be used with a second data visualization, which is obvious that a second data visualization is a second subset of data points. Keahey also discloses data aggregation/binning or filtering that can be applied. This teaches the claimed subject matter of from the statistical data distribution, a second subset of data points from the plurality of data points and a filtering or grouping operation on each data point of the second subset.
The motivation for the above is to have accurate statistical data distribution for better operation of subset of data points.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur by applying a machine learning model to determine, from the statistical data distribution, a second subset of data points from the plurality of data points and performing a filtering or grouping operation on each data point of the second subset of data points as taught by Keahey.
Regarding claim 17, Setlur teaches “The method of claim 1, wherein the first subset of data points is selected further based on a chart type of the data visualization.” (may determine the data visualization type for the data visualization based, at least in part, on the determined user intent.... The data visualization type may be one of: a Bar Chart, a Line Chart, a Scatter Plot, a Pie Chart, a Map, or a Text Table; Col 12, Line11-16);
Setlur discloses determining a visualization type based on intent which teaches the selection of chart type of the data visualization.
Regarding claim 18, Setlur teaches “The method of claim 1, wherein the data visualization is a Sankey chart, a tree map, a stacked bar graph, a scatter plot, or a line chart.” (The data visualization type may be one of: a Bar Chart, a Line Chart, a Scatter Plot, a Pie Chart, a Map, or a Text Table; Col 12, Line 14-16);
Setlur discloses at least a bar chart, scatter plot and line chart which teaches the claimed subject matter of at least a stacked bar graph, a scatter plot, or a line chart.
Regarding claim 19, Setlur teaches “A computing device executing a browser application, comprising:” (computing device (e.g., the computing device 200); Col 20, Line 6-7); (generating data visualizations; Col 4, Line 17-18); (executed by a web browser on a user's computing device; Col 9, Line 34-35);
“a display;” (having a display (e.g., the display 212); Col 20, Line 7-8);
“one or more processors; and” (one or more processors (e.g., the processors 202); Col 20, Line 8-9);
“memory coupled to the one or more processors, the memory storing one or more programs configured for execution by the one or more processors, the one or more programs including instructions for:” (memory (e.g., the memory 206) storing (806) one or more programs configured for execution by the one or more processors; Col 20, Line 9-11);
Claim 19 is directed to a computing device and its limitations are similar in scope and functions performed by the method of claim 1. Therefore, claim 19 limitations are also rejected with the same rationale as regarding claim 1.
Regarding claim 20, Setlur teaches “A non-transitory computer-readable storage medium storing one or more programs configured for execution by one or more processors of a computing device executing a browser application, the one or more programs comprising instructions for:” (a non-transitory computer-readable storage medium stores one or more programs configured for execution by a computing device; Col 3, Line 36-38); (executed by a web browser on a user's computing device; Col 9, Line 34-35);
Claim 20 is directed to a non-transitory computer-readable storage medium and its limitations are similar in scope and functions performed by the method of claim 1. Therefore, claim 20 limitations are also rejected with the same rationale as regarding claim 1.
Claim(s) 4-5 are rejected under 35 U.S.C. 103 as being unpatentable over SETLUR in view of KEAHEY and in further view of JAIN and in further view of GAVVALA (No. US-20240103848-A1 “Gavvala”).
Regarding claim 4, while Setlur, Keahey and Jain do not teach the limitation, Gavvala teaches “The method of claim 2, further comprising:
generating a distinct computation pipeline for each data segment, of the multiple data segments, to independently process the data segment.” (computational tasks may be delegated to a specific processing component ... hardware accelerators 185 a-n may include ... GPUs; Para 0021);
Gavvala discloses a distinct processing pipeline which is firmware focused. This supports computation pipeline for processing different data segments which teach the claimed subject matter.
The motivation for the above is to have an efficient pipeline of computation for easier processing of data visualization.
It would be obvious to a person skilled in the art that when generate data visualization there are computational tasks delegated to generate the visualization.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur, Keahey and Jain by generating a distinct computation pipeline for each data segment, of the multiple data segments, to independently process the data segment as taught by Gavvala.
Regarding claim 5, Setlur further teaches “for each data region:
determining a value for a visual change parameter for the data visualization when data values of the data region are included an existing rendering of the data visualization; and” (The visual variables include information that encode how the data visualization will look (e.g., data visualization type, what data points will be displayed or represented as visual marks, the color scheme of visual marks, or emphasizing certain visual marks); Col 11, Line 33-37);
While Setlur does not teach “at a respective computation pipeline corresponding to a respective data segment: dividing the respective data segment into one or more data regions”, “in accordance with a determination that the value for visual change parameter satisfies a threshold value” and “adding the data region to the at least a portion of each data segment; and reapplying the first algorithm to the a least a portion of each data segment”.
Keahey teaches “dividing the respective data segment into one or more data regions; and” (the data visualizations 144 may make up a portion of the dataset(s) 142. ... organized into sets for evaluation by the visualization evaluation module 124; Para 0041);
“in accordance with a determination that the value for visual change parameter satisfies a threshold value:” (selections to specify such criteria (e.g., parameters) to be used to determine the candidate data visualizations; Para 0064); (the data analytics application 122 can select each of the candidate data visualizations that have a score exceeding a threshold value; Para 0065);
“adding the data region to the at least a portion of each data segment; and
reapplying the first algorithm to the a least a portion of each data segment.” (the data analytics application 122 can select and/or generate additional data visualizations and again iterate the above processes to select and/or generate additional data visualizations and add them to the collection. In doing so, the visualization evaluation module 124 can re-compute the transition; Para 0068);
Setlur discloses determining visual parameters through visual variables for data visualization. In combination with Keahey disclosing scores and the threshold value, teaching the claimed subject matter.
The motivation for the above is to have an efficient pipeline of computation for easier processing of data visualization.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur by dividing the respective data segment into one or more data regions, in accordance with a determination that the value for visual change parameter satisfies a threshold value, adding the data region to the at least a portion of each data segment and reapplying the first algorithm to the a least a portion of each data segment as taught by Keahey.
Selfur, Keahey and Jain fail to specifically disclose “at a respective computation pipeline corresponding to a respective data segment”.
Gavvala teaches “The method of claim 4, further comprising:
at a respective computation pipeline corresponding to a respective data segment:” (computational tasks may be delegated to a specific processing component ... hardware accelerators 185 a-n may include ... GPUs; Para 0021);
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur, Keahy and Jain by a respective computation pipeline corresponding to a respective data segment as taught by Gavvala.
The motivation for the above is to have an efficient pipeline of computation for easier processing of data visualization.
Claim(s) 6 and 8-11 are rejected under 35 U.S.C. 103 as being unpatentable over SETLUR in view of KEAHEY and JAIN and in further view of UDESHI (No. US-20140108419-A1 “Udeshi”).
Regarding claim 6, while Selfur, Keahey and Jain do not teach the limitations, Udeshi teaches “The method of claim 1, further comprising:
after obtaining the dataset:
generating a data structure that includes a plurality of nodes; and” (Elevation data may be stored in a database characterized by a quadtree structure.... composed of nodes and leaves; Para 0023);
“assigning each data point of the dataset to a respective node of the data structure according to a spatial location of the respective data point in the data visualization.” (a quadtree containing elevation values stored in a database may contain two or more types of rows. ... position of the point or area covered may be expressed as latitude and longitude coordinates; Para 0050);
Udeshi discloses a quadtree which is assigned a node based on the position expressed by the latitude and longitude coordinates which teaches the claimed subject matter of spatial location of the respective data points. In combination with what is taught by Jain, Keahey and Setlur with data visualization teach the claimed subject matter.
Selfur, Keahey, Jain and Udeshi are analogous art as they are related to data structure.
The motivation for the above is to have an efficient data structure for the dataset of the visualization.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur, Jain and Keahey by generating a data structure that includes a plurality of nodes and assigning each data point of the dataset to a respective node of the data structure according to a spatial location of the respective data point in the data visualization as taught by Udeshi.
Regarding claim 8, while Selfur, Keahey and Jain do not teach the limitations, Udeshi teaches “The method of claim 6, wherein the data structure comprises a quadtree data structure.” (Elevation data may be stored in a database characterized by a quadtree structure; Para 0023);
The motivation for the above is to have an efficient data structure like quad tree for accurate data visualization.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Selfur, Keahey and Jain by the data structure comprises a quadtree data structure as taught by Udeshi.
Regarding claim 9, while Selfur, Keahey and Jain do not teach the limitations, Udeshi teaches “The method of claim 6, wherein:
the data visualization occupies a spatial area; and” (Quadtrees may be used in conjunction with a two-dimensional map; Para 0025);
“the method includes:
partitioning the spatial area into four quadrants; and” (divide a two-dimensional map into four partitions; Para 0025);
“for a respective quadrant:
recursively partitioning the quadrant to sub-quadrants in accordance with a determination that a first set of criteria is satisfied; and” (sub-divide the created partitions further depending on the data contained; Para 0025);
assigning a respective data point to a respective sub-quadrant according to respective coordinates of the data point.” (sub-divide the created partitions further depending on the data contained; Para 0025); (determining the various smaller nodes covering the desired node; Para 0029);
Udeshi discloses a 2D map the showcases the spatial area occupying. The 2D map also has four partitions that teach the spatial area into four quadrants. In addition, Udeshi discloses dividing the partitions further depending on data contained which can showcase recursive partitioning and criteria being satisfied.
The motivation for the above is to accurate visualization of spatial area for easier overall data visualization.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Selfur, Keahey and Jain by the data visualization occupies a spatial area, partitioning the spatial area into four quadrants, for a respective quadrant: recursively partitioning the quadrant to sub-quadrants in accordance with a determination that a first set of criteria is satisfied and assigning a respective data point to a respective sub-quadrant according to respective coordinates of the data point as taught by Udeshi.
Regarding claim 10, while Setlur does not teach the limitation, Keahey teaches “The method of claim 9, wherein the first set of criteria includes a criterion that a number of data points corresponding to the respective quadrant exceeds a threshold number of data points.” (select each of the candidate data visualizations that have a score exceeding a threshold value; Para 0065);
Keahey discloses candidate data visualizations that can exceed the threshold value which relates to the criteria that a number of data points exceed a threshold value. Keahey in combination with the quadrants of the quad tree teach the overall claimed subject matter.
It would be obvious to a person skilled in the art to have criteria in relation to the data points because when rendering a visualization, factors like criteria of the number of data points contributes to the generation of the data visualization.
Thus, it would have been obvious to one of ordinary skill in the art at the effective filing date of the claimed invention to modify Setlur by wherein the first set of criteria includes a criterion that a number of data points corresponding to the respective quadrant exceeds a threshold number of data points as taught by Keahey.
Regarding claim 11, Setlur teaches “in accordance with a determination that the first node includes one or more data points that are excluded from the final subset of data points:
re-rendering the first region of the data visualization to include one or more additional data marks, corresponding to the one or more data points; and
displaying the re-rendered first region of the data visualization.” (generates and displays a data visualization (or an updated data visualization) of retrieved data sets; Col 4, Line 36-38); (generates and displays a data visualization, including a plurality of visual marks representing data retrieved from the data source; Col 2, Line 40-42);
Setlur does not teach “The method of claim 6, further comprising:
after displaying, on the browser application, the data visualization:
receiving user selection of a first region of the data visualization, the first region including at least one data mark of the plurality of data marks; and in response to receiving the user selection of the first region of the data visualization:
identifying a first node, in the data structure, corresponding to the first region of the data visualization;”
Udeshi teaches “The method of claim 6, further comprising:
after displaying, on the browser application, the data visualization:
receiving user selection of a first region of the data visualization, the first region including at least one data mark of the plurality of data marks; and” (Given a user-selected area, such as a rectangle, index tiles that overlap the points covered by the user-selected area may be determined; Para 0090);
“in response to receiving the user selection of the first region of the data visualization:
identifying a first node, in the data structure, corresponding to the first region of the data visualization;” (index tiles that overlap the points covered by the user-selected area may be determined; Para 0090);
Udeshi discloses user selected area that has point that showcase the at least one data mark of the plurality of data marks and index tiles which relates to the nodes in the data structure. Setlur discloses generating and displaying the data visualization including a plurality of marks, thus teaching the claimed subject matter. In combination they teach the final displaying of the data visualization.
The motivation for the above is to have an accurate display of data visualization for user friendly viewing.
It would be obvious to a person skilled in the art that the user selected area in a data visualization can improve the final generation of the data visualization.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur, Keahey and Jain by after displaying, on the browser application, the data visualization: receiving user selection of a first region of the data visualization, the first region including at least one data mark of the plurality of data marks and in response to receiving the user selection of the first region of the data visualization: identifying a first node, in the data structure, corresponding to the first region of the data visualization as taught by Udeshi.
Claim(s) 7 is rejected under 35 U.S.C. 103 as being unpatentable over SETLUR in view of KEAHEY and JAIN and in further view of UDESHI and ABADI (No. US-20200104324-A1 “Abadi”).
Regarding claim 7, while Selfur, Keahey, Jain and Udeshi do not teach the limitation, Abadi teaches “The method of claim 6, further comprising storing each data point of the dataset in a binary data format in the data structure.” (The tree can be written in a binary format suitable for efficient loading into memory; Para 0013);
Abadi discloses binary format that loads into memory which teaches the claimed subject matter of storing data in a binary format. In combination with Udeshi that teaches quadtree, a tree written in binary can also appl to a quad tree.
Udeshi and Abadi are analogous art as they are related to data structure and tree formation.
The motivation for the above is to have an efficient format when storing tree nodes for accurate data visualization.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur, Keahey, Jain and Udeshi by storing each data point of the dataset in a binary data format in the data structure as taught by Abadi.
Claim(s) 12 is rejected under 35 U.S.C. 103 as being unpatentable over SETLUR in view of KEAHEY in view of Jain and in further view of FAN (No. CN-104966172-A “Fan”).
Regarding claim 12, while Setlur, Keahey and Jain do not teach the limitations, Fan teaches “The method of claim 1, further comprising:
after obtaining the dataset and prior to selecting the first subset of data points:
performing initial data cleaning and transformation.” (perform data transformation and cleaning; Pg.2, Para 16); (The data cleaning module performs data cleaning on the data to be cleaned.... the data normalization processing module normalizes the cleaned data and saves the data to the data pool; Pg.3, Para 2);
Fan discloses data cleaning and transformation. The cleaning and transformation are applied to data which can showcase thar after obtaining data, there is data cleaning and transformation. This in combination with Setlur, Keahey and Jain teach the claimed subject matter.
Setlur, Keahey, Jain and Fan are analogous art as they are related to data visualization.
The motivation for the above is to have efficient data processing for overall accurate data visualization.
Therefore, it would have been obvious for an ordinary skilled person in the art before the effective filing date of claimed invention to have modified Setlur, Keahey and Jain by performing initial data cleaning and transformation as taught by Fan.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US-20030079005-A1 (Myers) – Discloses optimizing a first route based on a first characteristic and optimizing a second route based on a second characteristic.
US-12013873-B2 (Sherman) – Discloses generating a data visualization by executing a data visualization data flow graph comprising a directed graph having a plurality of nodes. Each of the nodes specifies either a data retrieval operation or a data transformation operation and the data visualization comprise visual marks having a first set of characteristics, including a first mark type and one or more first visual mark encodings. A user specifies a second mark type and/or one or more second visual mark encodings.
US-9710430-B2 (Werner) – Discloses providing visual bundlers that group and represent specified data subsets of very large datasets in a manner that is expressive and intuitive for a user, and which provide a dynamic, configurable visualization that may be leveraged by the user to search, aggregate, or otherwise interact with the data of a very large dataset.
US-9978114-B2 (Subramaniyan) – Discloses a system for optimizing processing and display of large datasets is provided. The system includes a graphics processing and optimization (GPO) computing device. The GPO computing device is configured to store a dataset including a data point in a memory device, select the data point to display on a display device based on a first display request signal received via a user interface, and accelerate graphical processing of the dataset using optimization algorithms.
Yuan, J., Xiang, S., Xia, J., Yu, L., & Liu, S. (2020). Evaluation of sampling methods for scatterplots. IEEE Transactions on Visualization and Computer Graphics, 27(2), 1720-1730. (Year: 2020) – Discloses understanding the capability of sampling methods in preserving the density, outliers, and overall shape of a scatterplot.
Pham, V., Nguyen, N. V., & Dang, T. (2020, July). Scagcnn: Estimating visual characterizations of 2d scatterplots via convolution neural network. In Proceedings of the 11th international conference on advances in information technology (pp. 1-9). (Year: 2020) – Discloses a set of visual features that characterizes the data distribution of a 2D scatterplot and has been used in a wide range of applications.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to BRIGITER D PROTAZI whose telephone number is (571)272-7995. The examiner can normally be reached Monday - Friday 7:30-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Said A Broome can be reached at 5712722931. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/B.D.P./Examiner, Art Unit 2612
/ALICIA M HARRINGTON/Supervisory Patent Examiner, Art Unit 2615