DETAILED ACTION
This action is in response to the amendments and remarks filed 06/15/2026. Claims 1-5, 9-13, 16-17, and 19-23 are pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/15/2026 has been entered.
Claim Interpretation
Claim 16 refers to, “a computer readable storage medium”. Paragraph [0069] of the instant Specification states, “A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire”. Accordingly, the computer readable storage media is not interpreted to include transitory signals per se.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim(s) 1-5 are rejected under 35 U.S.C. 103 as being unpatentable over Mandal (A Topological Data Analysis Approach on Predicting Phenotypes from Gene Expression Data, 2020, AlCoB 2020, LNBI 12099, pp. 178-187) in view of Chazal (Subsampling Methods for Persistent Homology, published 2014, arXiv:1406.1901v1), and further in view of Bubenik (STATISTICAL TOPOLOGICAL DATA ANALYSIS USING PERSISTENCE LANDSCAPES, published 1/23/2015, arXiv:1207.6437v4), and Modarres (Hierarchical Floorplanner, published 4/17/1990, Patent no. 4,918,614).
Regarding claim 1, Mandal teaches [a] computer-implemented method of training a neural network for disease detection in a sample, comprising:
creating a training set based on topological summaries: “we use the gene expressions of subjects with and without Parkinson’s disease to generate topological summaries per subject. These summaries essentially act as unique fingerprints that describe the topology of the gene expression in a sample. We use these fingerprints to enhance the feature vector that is used for disease phenotype prediction, “ (Mandal, page 179, paragraph 5); “For each subject, we obtained a feature vector of 19,581 gene expression measurements (see Sect. 2.6) and a known class label 0 (161 control subjects) or 1 (264 affected subjects) according to the Parkinson’s disease phenotype. We split the data 80-20 into training and test sets, over 50 iterations, except for the computationally more intensive TDA-CNN where we considered 4 iterations after observing the results between iterations were nearly identical” (Mandal, page 184, paragraph 1).
… wherein the topological summaries are associated with different phenotypes: “The topological summaries of the weighted point clouds (persistence landscapes) are then used to construct a machine learning model to predict the phenotype (healthy or PD) for each sample” (Mandal, page 180, paragraph 2).
…creating the training set including at least:
receiving gene expression data associated with a plurality of subjects: “For each subject, we obtained a feature vector of 19,581 gene expression measurements (see Sect. 2.6) and a known class label 0 (161 control subjects) (plurality of subjects) or 1 (264 affected subjects) (plurality of subjects) according to the Parkinson’s disease phenotype” (Mandal, page 184, paragraph 1).
determining pair-wise similarities between genes in the gene expression data of all of the plurality of subjects: “We work under the hypothesis that the set X of all subjects’ samples (gene expression data of all the plurality of subjects), each encoded as a collection of gene expression values, can provide us with enough topological information to discern between healthy subjects and subjects with Parkinson’s disease. We denote by X a matrix of size
n
r
o
w
s
×
n
c
o
l
s
where each row corresponds to a subject and each column corresponds to a gene. Each entry
X
i
,
j
then corresponds to the j-th gene expression of the i-th subject.” (Mandal, page 179, paragraph 6); “Co-expression can be examined by computing pairwise correlations between gene expression measurements. Therefore, we construct a new matrix
X
-
from X, consisting of all pairwise distance correlations between genes” (Mandal, page 180, paragraph 2).
transforming the gene expression data into the topological summaries based on the pair-wise similarities: “Later, we show how to use the theory of persistent homology to determine the persistent topological landscapes present in the gene expression data of a sample, by first transforming it into a weighted point cloud. We do this transformation by utilizing the gene correlations (pair-wise similarities) across all available samples (matrix
X
-
). The topological summaries of the weighted point clouds (persistence landscapes) are then used to construct a machine learning model to predict the phenotype (healthy or PD) for each sample” (Mandal, page 180, paragraph 2).
the topological summaries including a persistence landscape for each degree of at least three homology degrees per subject of the plurality of subjects: “For each of the simplicial complexes we obtained persistence landscapes [5] for homology dimensions 0 and 1. Such landscapes are, for each homology degree, sequences {
λ
k
} of decreasing piecewise linear (PL) functions
λ
k
:
R
→
R
… After computation of all landscapes, for each subject we then obtained its average landscape” (Mandal, page 183, paragraph 3);
the transforming further including subsampling data points of the gene expression data by randomly selecting a subset of a preconfigured size, the sub-sampling preserving a topology associated with the topological summaries: “To mitigate the computational cost of our setup we used a subsampling approach, as studied in [11], so that instead of working with the entire set of genes at all times, for each subject we repeatedly subsampled smaller sets of
n
s
u
b
s
a
m
p
l
e
genes, obtaining several filtered simplicial complexes” (Mandal, page 183, paragraph 2).
training a neural network using the training set created based on the topological summaries: “Ultimately, we use the gene expressions of subjects with and without Parkinson’s disease to generate topological summaries per subject. These summaries essentially act as unique fingerprints that describe the topology of the gene expression in a sample. We use these fingerprints to enhance the feature vector that is used for disease phenotype prediction” (Mandal, page 179, paragraph 5); “For each subject, we obtained a feature vector of 19,581 gene expression measurements (see Sect. 2.6) and a known class label 0 (161 control subjects) or 1 (264 affected subjects) according to the Parkinson’s disease phenotype. We split the data 80-20 into training and test sets, over 50 iterations, except for the computationally more intensive TDA-CNN where we considered 4 iterations after observing the results between iterations were nearly identical” (Mandal, page 184, paragraph 1).
by feeding the persistence landscape for each degree of the plurality of homology degrees into the neural network: “In the TDA-CNN approach, for a given resolution ry, we fed each subject’s vectorized persistence landscape as a tensor of shape (rx, ry, 2), one channel per homology degree, into a Convolutional Neural Network (CNN)” (Mandal, page 184, paragraph 3)
… wherein the topological summaries are converted to a tensor and the tensor is fed into the neural network for training the neural network:
“We then quantized the resulting landscapes (topological summaries) by sampling
r
x
values evenly in the interval [0,
t
m
a
x
], where
t
m
a
x
is a value estimated from the data that corresponds to the last time of the filtration where there were changes in the persistent homology of the complex being processed. This results, for each persistence landscape, in a 2D array (tensor) of size
r
x
×
n
λ
” (Mandal, page 183, paragraph 4). As stated in paragraph [0022] of the instant Specification, a persistence landscape is an instance / example of a topological summary.
“For each subject, we obtained a feature vector of 19,581 gene expression measurements (see Sect. 2.6) and a known class label 0 (161 control subjects) or 1 (264 affected subjects) according to the Parkinson’s disease phenotype. We split the data 80-20 into training and test sets, over 50 iterations, except for the computationally more intensive TDA-CNN where we considered 4 iterations after observing the results between iterations were nearly identical” (Mandal, page 184, paragraph 1); “In the TDA-CNN approach, for a given resolution
r
y
, we fed each subject’s vectorized persistence landscape as a tensor of shape (
r
x
,
r
y
, 2), one channel per homology degree, into a Convolutional Neural Network (CNN)” (Mandal, page 184, paragraph 3).
… wherein the persistence landscape for said each degree includes a collection of functions fi(x) = y from real numbers to real numbers: “For each of the simplicial complexes we obtained persistence landscapes [5] for homology dimensions 0 and 1. Such landscapes are, for each homology degree, sequences {
λ
k
} of decreasing piecewise linear (PL) functions
λ
k
:
R
→
R
” (Mandal, page 183, paragraph 3)
… wherein the tensor has a shape of I times X times Y times D, wherein I is a total number of functions fi, X is a total number of x values, Y is a total number of y values and D is a total number of the plurality of homology degrees considered: “In the TDA-CNN approach, for a given resolution
r
y
, we fed each subject’s vectorized persistence landscape as a tensor of shape (
r
x
,
r
y
, 2), one channel per homology degree, into a Convolutional Neural Network (CNN)” (Mandal, page 184, paragraph 3).
… the neural network being trained to predict a phenotype of a sample associated with a subject whose phenotype is unknown: “The topological summaries of the weighted point clouds (persistence landscapes) are then used to construct a machine learning model to predict the phenotype (healthy or PD) for each sample.” (Mandal, page 180, paragraph 3)
Mandal relates to using topological data analysis to diagnose diseased phenotypes with machine learning and is analogous to the claimed invention.
While Mandal fails to disclose the further limitations of the claim, Chazal discloses a method of subsampling data points of the gene expression data by randomly selecting a subset of a preconfigured size, the sub-sampling preserving a topology associated with the topological summaries:
“For any positive integer m, let
X
=
{
x
1
,
…
,
x
m
}
⊂
X
be a sample of m points from the measure
μ
∈
P
(
X
)
. The corresponding persistence landscape is
λ
X
” (Chazal, page 4, paragraph 6)
PNG
media_image1.png
321
879
media_image1.png
Greyscale
”Left: 3D shapes of the first experiment. Middle and Left: 500 random points from the magnetometer data of the second experiment.” (Chazal, page 7, Figure 3)
“In practice, each shape consists of a 3D point cloud embedded in the Euclidean space, with a number of vertices that ranges from 7K to 40K … For n = 100 times we subsample m = 300 points (preconfigured size) from each shape; then we select the closest subsample to the corresponding original point cloud and compute
4
×
n
persistence diagrams (dimension 1), one for each subsample” (Chazal, page 7, paragraph 3)
“We study the risk of two estimators and we prove that the subsampling approach carries stable topological information (preserving a topology) while achieving a great reduction in computational complexity” (Chazal, page 1, Abstract)
Chazal relates to subsampling point clouds of data to reduce the cost of persistent homology analysis and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the primary reference to randomly subsample data such that subsamples maintain topological properties, as disclosed by Chazal. Chazal demonstrated that topological data analysis through these subsets accurately approximates persistent homology of the full set of data points, while being computationally faster and simple to execute. This is particularly useful when trying to perform topological data analysis is prohibitively expensive due to large data sets. See Chazal, page 1, Abstract & pages 7-9.
While Chazal fails to disclose the further limitations of the claim, Bubenik discloses a method, wherein the topological summaries including a persistence landscape for each degree of at least three homology degrees per subject of the plurality of subjects: “We sample 100 points from the uniform distribution on the unit cube
[
0,1
]
3
, and calculate the persistence landscapes in degrees 0, 1 and 2 (three homology degrees) of the corresponding Vietoris-Rips complex” (Bubenik, page 13, paragraph 4)
Bubenik relates to topological data analysis with machine learning and is analogous to the claimed invention. The existing combination teaches a method of training neural networks with persistence landscapes. Bubenik teaches a method of calculating persistence landscapes for homology degrees up to three. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the existing combination with Bubenik by calculating topological summaries for three homology degrees instead of two. This would achieve the predictable result of calculating persistence landscapes that includes more homologies from different degrees, with the existing combination’s method of calculating and applying topological data summaries and Bubenik’s method of using three degrees of homology in calculating persistence landscapes performing the same together as they did separately. (MPEP 2143 I. (A) Combining prior art elements according to known methods to yield predictable results).
While Bubenik fails to disclose the further limitations of the claim, Modarres discloses a method, wherein the tensor has a shape of I times X times Y times D, wherein I is a total number of functions fi: “The Net Table is a two-dimensional array (tensor) which contains a number of columns equal to the total number of functions through the hierarchy, and a number of rows equal to the number of nets specified by the user” (Modarres, column 11, paragraph 7)
Modarres is reasonably pertinent to the representation of functions in tensors and is analogous to the claimed invention. The existing combination teaches a method of representing persistence landscapes as three-dimensional tensors split along X values, Y values, and homology degrees. Modarres teaches tensors with a dimension for functions. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the existing combination and Modarres by additionally splitting persistence landscape tensors along a fourth dimension for landscape functions. This would achieve the predictable result of explicitly splitting the data along the persistence landscape function values, with the existing combination’s persistence landscape calculations and Modarres’ feature dimension performing the same together as they did separately. (MPEP 2143 I. (A) Combining prior art elements according to known methods to yield predictable results).
Regarding claim 2, the rejection of claim 1 is incorporated. Mandal further teaches a method of receiving a new sample; creating a new topological summary associated with the new sample based on the pairwise similarities; and inputting the new topological summary to the neural network, the neural network predicting the new sample’s phenotype: “Later, we show how to use the theory of persistent homology to determine the persistent topological landscapes present in the gene expression data of a sample (new sample), by first transforming it into a weighted point cloud ... The topological summaries (of the weighted point clouds (persistence landscapes) are then used to construct a machine learning model to predict the phenotype (healthy or PD) for each sample” (Mandal, page 180, paragraph 3); “In the TDA-CNN approach, for a given resolution
r
y
, we fed each subject’s vectorized persistence landscape ... into a Convolutional Neural Network (CNN)” (Mandal, page 184, paragraph 3).
Regarding claim 3, the rejection of claim 1 is incorporated. Mandal further teaches a method, wherein the neural network includes a convolutional neural network: “In the TDA-CNN approach, for a given resolution
r
y
, we fed each subject’s vectorized persistence landscape as a tensor of shape (
r
x
,
r
y
, 2), one channel per homology degree, into a Convolutional Neural Network (CNN)” (Mandal, page 184, paragraph 3).
Regarding claim 4, the rejection of claim 1 is incorporated. Mandal further teaches a method, wherein the pair-wise similarities include distance measures between pairs of genes in the gene expression data: “Co-expression can be examined by computing pairwise correlations between gene expression measurements. Therefore, we construct a new matrix
X
-
from X, consisting of all pairwise distance correlations between genes” (Mandal, page 180, paragraph 2).
Regarding claim 5, the rejection of claim 1 is incorporated. Mandal further teaches a method, wherein the pair-wise similarities are used to create a point cloud, the point cloud used to transform the gene expression data into the topological summaries: “Later, we show how to use the theory of persistent homology to determine the persistent topological landscapes present in the gene expression data of a sample, by first transforming it into a weighted point cloud. We do this transformation by utilizing the gene correlations (pair-wise similarities) across all available samples (matrix
X
-
). The topological summaries of the weighted point clouds (persistence landscapes) are then used to construct a machine learning model to predict the phenotype (healthy or PD) for each sample” (Mandal, page 180, paragraph 2).
Claim(s) 9-13, 16-17, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Mandal (A Topological Data Analysis Approach on Predicting Phenotypes from Gene Expression Data, 2020, AlCoB 2020, LNBI 12099, pp. 178-187) in view of Chazal (Subsampling Methods for Persistent Homology, published 2014, arXiv:1406.1901v1), and further in view of Bubenik (STATISTICAL TOPOLOGICAL DATA ANALYSIS USING PERSISTENCE LANDSCAPES, published 1/23/2015, arXiv:1207.6437v4), Modarres (Hierarchical Floorplanner, published 4/17/1990, Patent no. 4,918,614), and Brandsma (Predicting and Addressing Severe Disease in Individuals with Sepsis, PCT filed 12/24/2020, US 2023/0018537 A1).
Regarding claim 9, Mandal teaches instructions to:
create a training set based on topological summaries: “we use the gene expressions of subjects with and without Parkinson’s disease to generate topological summaries per subject. These summaries essentially act as unique fingerprints that describe the topology of the gene expression in a sample. We use these fingerprints to enhance the feature vector that is used for disease phenotype prediction, “ (Mandal, page 179, paragraph 5); “For each subject, we obtained a feature vector of 19,581 gene expression measurements (see Sect. 2.6) and a known class label 0 (161 control subjects) or 1 (264 affected subjects) according to the Parkinson’s disease phenotype. We split the data 80-20 into training and test sets, over 50 iterations, except for the computationally more intensive TDA-CNN where we considered 4 iterations after observing the results between iterations were nearly identical” (Mandal, page 184, paragraph 1).
… wherein the topological summaries are associated with different phenotypes: “The topological summaries of the weighted point clouds (persistence landscapes) are then used to construct a machine learning model to predict the phenotype (healthy or PD) for each sample” (Mandal, page 180, paragraph 2).
… the processor configured to create the training set at least by performing operations of:
receive gene expression data associated with a plurality of subjects: “For each subject, we obtained a feature vector of 19,581 gene expression measurements (see Sect. 2.6) and a known class label 0 (161 control subjects) (plurality of subjects) or 1 (264 affected subjects) (plurality of subjects) according to the Parkinson’s disease phenotype” (Mandal, page 184, paragraph 1).
determine pair-wise similarities between genes in the gene expression data of all of the plurality of subjects: “We work under the hypothesis that the set X of all subjects’ samples (gene expression data of all the plurality of subjects), each encoded as a collection of gene expression values, can provide us with enough topological information to discern between healthy subjects and subjects with Parkinson’s disease. We denote by X a matrix of size
n
r
o
w
s
×
n
c
o
l
s
where each row corresponds to a subject and each column corresponds to a gene. Each entry
X
i
,
j
then corresponds to the j-th gene expression of the i-th subject.” (Mandal, page 179, paragraph 6); “Co-expression can be examined by computing pairwise correlations between gene expression measurements. Therefore, we construct a new matrix
X
-
from X, consisting of all pairwise distance correlations between genes” (Mandal, page 180, paragraph 2).
transform the gene expression data into the topological summaries based on the pair-wise similarities: “Later, we show how to use the theory of persistent homology to determine the persistent topological landscapes present in the gene expression data of a sample, by first transforming it into a weighted point cloud. We do this transformation by utilizing the gene correlations (pair-wise similarities) across all available samples (matrix
X
-
). The topological summaries of the weighted point clouds (persistence landscapes) are then used to construct a machine learning model to predict the phenotype (healthy or PD) for each sample” (Mandal, page 180, paragraph 2).
the topological summaries including a persistence landscape for each degree of at least three homology degrees per subject of the plurality of subjects: “For each of the simplicial complexes we obtained persistence landscapes [5] for homology dimensions 0 and 1. Such landscapes are, for each homology degree, sequences {
λ
k
} of decreasing piecewise linear (PL) functions
λ
k
:
R
→
R
… After computation of all landscapes, for each subject we then obtained its average landscape” (Mandal, page 183, paragraph 3);
the transforming further including subsampling data points of the gene expression data by randomly selecting a subset of a preconfigured size, the sub-sampling preserving a topology associated with the topological summaries: “To mitigate the computational cost of our setup we used a subsampling approach, as studied in [11], so that instead of working with the entire set of genes at all times, for each subject we repeatedly subsampled smaller sets of
n
s
u
b
s
a
m
p
l
e
genes, obtaining several filtered simplicial complexes” (Mandal, page 183, paragraph 2).
train a neural network using the training set created based on the topological summaries: “Ultimately, we use the gene expressions of subjects with and without Parkinson’s disease to generate topological summaries per subject. These summaries essentially act as unique fingerprints that describe the topology of the gene expression in a sample. We use these fingerprints to enhance the feature vector that is used for disease phenotype prediction” (Mandal, page 179, paragraph 5); “For each subject, we obtained a feature vector of 19,581 gene expression measurements (see Sect. 2.6) and a known class label 0 (161 control subjects) or 1 (264 affected subjects) according to the Parkinson’s disease phenotype. We split the data 80-20 into training and test sets, over 50 iterations, except for the computationally more intensive TDA-CNN where we considered 4 iterations after observing the results between iterations were nearly identical” (Mandal, page 184, paragraph 1).
by feeding the persistence landscape for each degree of the plurality of homology degrees into the neural network: “In the TDA-CNN approach, for a given resolution ry, we fed each subject’s vectorized persistence landscape as a tensor of shape (rx, ry, 2), one channel per homology degree, into a Convolutional Neural Network (CNN)” (Mandal, page 184, paragraph 3)
… wherein the topological summaries are converted to a tensor and the tensor is fed into the neural network for training the neural network:
“We then quantized the resulting landscapes (topological summaries) by sampling
r
x
values evenly in the interval [0,
t
m
a
x
], where
t
m
a
x
is a value estimated from the data that corresponds to the last time of the filtration where there were changes in the persistent homology of the complex being processed. This results, for each persistence landscape, in a 2D array (tensor) of size
r
x
×
n
λ
” (Mandal, page 183, paragraph 4). As stated in paragraph [0022] of the instant Specification, a persistence landscape is an instance / example of a topological summary.
“For each subject, we obtained a feature vector of 19,581 gene expression measurements (see Sect. 2.6) and a known class label 0 (161 control subjects) or 1 (264 affected subjects) according to the Parkinson’s disease phenotype. We split the data 80-20 into training and test sets, over 50 iterations, except for the computationally more intensive TDA-CNN where we considered 4 iterations after observing the results between iterations were nearly identical” (Mandal, page 184, paragraph 1); “In the TDA-CNN approach, for a given resolution
r
y
, we fed each subject’s vectorized persistence landscape as a tensor of shape (
r
x
,
r
y
, 2), one channel per homology degree, into a Convolutional Neural Network (CNN)” (Mandal, page 184, paragraph 3).
… wherein the persistence landscape for said each degree includes a collection of functions fi(x) = y from real numbers to real numbers: “For each of the simplicial complexes we obtained persistence landscapes [5] for homology dimensions 0 and 1. Such landscapes are, for each homology degree, sequences {
λ
k
} of decreasing piecewise linear (PL) functions
λ
k
:
R
→
R
” (Mandal, page 183, paragraph 3)
… wherein the tensor has a shape of I times X times Y times D, wherein I is a total number of functions fi, X is a total number of x values, Y is a total number of y values and D is a total number of the plurality of homology degrees considered: “In the TDA-CNN approach, for a given resolution
r
y
, we fed each subject’s vectorized persistence landscape as a tensor of shape (
r
x
,
r
y
, 2), one channel per homology degree, into a Convolutional Neural Network (CNN)” (Mandal, page 184, paragraph 3).
… the neural network being trained to predict a phenotype of a sample associated with a subject whose phenotype is unknown: “The topological summaries of the weighted point clouds (persistence landscapes) are then used to construct a machine learning model to predict the phenotype (healthy or PD) for each sample.” (Mandal, page 180, paragraph 3)
Mandal relates to using topological data analysis to diagnose diseased phenotypes with machine learning and is analogous to the claimed invention.
While Mandal fails to disclose the further limitations of the claim, Chazal discloses a method of subsampling data points of the gene expression data by randomly selecting a subset of a preconfigured size, the sub-sampling preserving a topology associated with the topological summaries:
“For any positive integer m, let
X
=
{
x
1
,
…
,
x
m
}
⊂
X
be a sample of m points from the measure
μ
∈
P
(
X
)
. The corresponding persistence landscape is
PNG
media_image1.png
321
879
media_image1.png
Greyscale
” Left: 3D shapes of the first experiment. Middle and Left: 500 random points from the magnetometer data of the second experiment.” (Chazal, page 7, Figure 3)
“In practice, each shape consists of a 3D point cloud embedded in the Euclidean space, with a number of vertices that ranges from 7K to 40K … For n = 100 times we subsample m = 300 points (preconfigured size) from each shape; then we select the closest subsample to the corresponding original point cloud and compute
4
×
n
persistence diagrams (dimension 1), one for each subsample” (Chazal, page 7, paragraph 3)
“We study the risk of two estimators and we prove that the subsampling approach carries stable topological (preserving a topology) information while achieving a great reduction in computational complexity” (Chazal, page 1, Abstract)
Chazal relates to subsampling point clouds of data to reduce the cost of persistent homology analysis and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the primary reference to randomly subsample data such that subsamples maintain topological properties, as disclosed by Chazal. Chazal demonstrated that topological data analysis through these subsets accurately approximates persistent homology of the full set of data points, while being computationally faster and simple to execute. This is particularly useful when trying to perform topological data analysis is prohibitively expensive due to large data sets. See Chazal, page 1, Abstract & pages 7-9.
While Chazal fails to disclose the further limitations of the claim, Bubenik discloses a method, wherein the topological summaries including a persistence landscape for each degree of at least three homology degrees per subject of the plurality of subjects: “We sample 100 points from the uniform distribution on the unit cube [0, 1]^3, and calculate the persistence landscapes in degrees 0, 1 and 2 (homology degrees) of the corresponding Vietoris-Rips complex” (Bubenik, page 13, paragraph 4)
Bubenik relates to topological data analysis with machine learning and is analogous to the claimed invention. the existing combination teaches a method of training neural networks with persistence landscapes. Bubenik teaches a method of calculating persistence landscapes for homology degrees up to three. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the existing combination with Bubenik by using Bubenik’s method to calculate persistence landscapes. This would achieve the predictable result of calculating persistence landscapes for a finite number of homology degrees, with the existing combination’s method of training neural networks and Bubenik’s method of calculating persistence landscapes performing the same together as they did separately. (MPEP 2143 I. (A) Combining prior art elements according to known methods to yield predictable results).
While Bubenik fails to disclose the further limitations of the claim, Modarres discloses a method, wherein the tensor has a shape of I times X times Y times D, wherein I is a total number of functions fi: “The Net Table is a two-dimensional array (tensor) which contains a number of columns equal to the total number of functions through the hierarchy, and a number of rows equal to the number of nets specified by the user” (Modarres, column 11, paragraph 7)
Modarres is reasonably pertinent to the representation of functions in tensors and is analogous to the claimed invention. The existing combination teaches a method of representing persistence landscapes as three-dimensional tensors split along X values, Y values, and homology degrees. Modarres teaches tensors with a dimension for functions. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the existing combination and Modarres by additionally splitting persistence landscape tensors along a fourth dimension for landscape functions. This would achieve the predictable result of explicitly splitting the data along the persistence landscape function values, with the existing combination’s persistence landscape calculations and Modarres’ feature dimension performing the same together as they did separately. (MPEP 2143 I. (A) Combining prior art elements according to known methods to yield predictable results).
While Modarres fails to disclose the further limitations of the claim, Brandsma teaches [a] system comprising: a processor; and a memory device coupled with the processor; the processor configured to at least: “In embodiments, there are provided systems for predicting severe disease in an individual with sepsis or at risk of developing sepsis, comprising: one or more processors; a memory; a communication platform” (Brandsma, [0019]).
Brandsma relates to neural networks and topological data analysis for disease phenotype prediction and is analogous to the claimed invention. The existing combination teaches a computational method of predicting disease phenotypes. The claimed invention improves upon this method by executing its method on computer hardware. Brandsma teaches computer hardware capable of running neural network methods for predicting phenotypes, applicable to the existing combination. A person of ordinary skill in the art would have recognized that running the existing combination’s method on Brandsma’s hardware would lead to the predictable result of the computational method being executed as described, and would improve the known device by allowing the method to produce concrete results on a computer (MPEP 2143 I. (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results).
The analysis of claims 10-13 mirrors that of claims 2-5, with the exception that claims 10-13 are directed to generic computer hardware which executes the methods of claims 2-5. This generic hardware is taught by Brandsma, as discussed regarding claim 9. Thus, claims 10-13 are rejected under the same rationales used for claims 2-5, respectively.
Regarding claim 16, Mandal teaches instructions to:
create a training set based on topological summaries: “we use the gene expressions of subjects with and without Parkinson’s disease to generate topological summaries per subject. These summaries essentially act as unique fingerprints that describe the topology of the gene expression in a sample. We use these fingerprints to enhance the feature vector that is used for disease phenotype prediction, “ (Mandal, page 179, paragraph 5); “For each subject, we obtained a feature vector of 19,581 gene expression measurements (see Sect. 2.6) and a known class label 0 (161 control subjects) or 1 (264 affected subjects) according to the Parkinson’s disease phenotype. We split the data 80-20 into training and test sets, over 50 iterations, except for the computationally more intensive TDA-CNN where we considered 4 iterations after observing the results between iterations were nearly identical” (Mandal, page 184, paragraph 1).
… wherein the topological summaries are associated with different phenotypes: “The topological summaries of the weighted point clouds (persistence landscapes) are then used to construct a machine learning model to predict the phenotype (healthy or PD) for each sample” (Mandal, page 180, paragraph 2).
… the processor configured to create the training set at least by performing operations of:
receive gene expression data associated with a plurality of subjects: “For each subject, we obtained a feature vector of 19,581 gene expression measurements (see Sect. 2.6) and a known class label 0 (161 control subjects) (plurality of subjects) or 1 (264 affected subjects) (plurality of subjects) according to the Parkinson’s disease phenotype” (Mandal, page 184, paragraph 1).
determine pair-wise similarities between genes in the gene expression data of all of the plurality of subjects: “We work under the hypothesis that the set X of all subjects’ samples (gene expression data of all the plurality of subjects), each encoded as a collection of gene expression values, can provide us with enough topological information to discern between healthy subjects and subjects with Parkinson’s disease. We denote by X a matrix of size
n
r
o
w
s
×
n
c
o
l
s
where each row corresponds to a subject and each column corresponds to a gene. Each entry
X
i
,
j
then corresponds to the j-th gene expression of the i-th subject.” (Mandal, page 179, paragraph 6); “Co-expression can be examined by computing pairwise correlations between gene expression measurements. Therefore, we construct a new matrix
X
-
from X, consisting of all pairwise distance correlations between genes” (Mandal, page 180, paragraph 2).
transform the gene expression data into the topological summaries based on the pair-wise similarities: “Later, we show how to use the theory of persistent homology to determine the persistent topological landscapes present in the gene expression data of a sample, by first transforming it into a weighted point cloud. We do this transformation by utilizing the gene correlations (pair-wise similarities) across all available samples (matrix
X
-
). The topological summaries of the weighted point clouds (persistence landscapes) are then used to construct a machine learning model to predict the phenotype (healthy or PD) for each sample” (Mandal, page 180, paragraph 2).
the topological summaries including a persistence landscape for each degree of at least three homology degrees per subject of the plurality of subjects: “For each of the simplicial complexes we obtained persistence landscapes [5] for homology dimensions 0 and 1. Such landscapes are, for each homology degree, sequences {
λ
k
} of decreasing piecewise linear (PL) functions
λ
k
:
R
→
R
… After computation of all landscapes, for each subject we then obtained its average landscape” (Mandal, page 183, paragraph 3);
the transforming further including subsampling data points of the gene expression data by randomly selecting a subset of a preconfigured size, the sub-sampling preserving a topology associated with the topological summaries: “To mitigate the computational cost of our setup we used a subsampling approach, as studied in [11], so that instead of working with the entire set of genes at all times, for each subject we repeatedly subsampled smaller sets of
n
s
u
b
s
a
m
p
l
e
genes, obtaining several filtered simplicial complexes” (Mandal, page 183, paragraph 2).
train a neural network using the training set created based on the topological summaries: “Ultimately, we use the gene expressions of subjects with and without Parkinson’s disease to generate topological summaries per subject. These summaries essentially act as unique fingerprints that describe the topology of the gene expression in a sample. We use these fingerprints to enhance the feature vector that is used for disease phenotype prediction” (Mandal, page 179, paragraph 5); “For each subject, we obtained a feature vector of 19,581 gene expression measurements (see Sect. 2.6) and a known class label 0 (161 control subjects) or 1 (264 affected subjects) according to the Parkinson’s disease phenotype. We split the data 80-20 into training and test sets, over 50 iterations, except for the computationally more intensive TDA-CNN where we considered 4 iterations after observing the results between iterations were nearly identical” (Mandal, page 184, paragraph 1).
by feeding the persistence landscape for each degree of the plurality of homology degrees into the neural network: “In the TDA-CNN approach, for a given resolution ry, we fed each subject’s vectorized persistence landscape as a tensor of shape (rx, ry, 2), one channel per homology degree, into a Convolutional Neural Network (CNN)” (Mandal, page 184, paragraph 3)
… wherein the topological summaries are converted to a tensor and the tensor is fed into the neural network for training the neural network:
“We then quantized the resulting landscapes (topological summaries) by sampling
r
x
values evenly in the interval [0,
t
m
a
x
], where
t
m
a
x
is a value estimated from the data that corresponds to the last time of the filtration where there were changes in the persistent homology of the complex being processed. This results, for each persistence landscape, in a 2D array (tensor) of size
r
x
×
n
λ
” (Mandal, page 183, paragraph 4). As stated in paragraph [0022] of the instant Specification, a persistence landscape is an instance / example of a topological summary.
“For each subject, we obtained a feature vector of 19,581 gene expression measurements (see Sect. 2.6) and a known class label 0 (161 control subjects) or 1 (264 affected subjects) according to the Parkinson’s disease phenotype. We split the data 80-20 into training and test sets, over 50 iterations, except for the computationally more intensive TDA-CNN where we considered 4 iterations after observing the results between iterations were nearly identical” (Mandal, page 184, paragraph 1); “In the TDA-CNN approach, for a given resolution
r
y
, we fed each subject’s vectorized persistence landscape as a tensor of shape (
r
x
,
r
y
, 2), one channel per homology degree, into a Convolutional Neural Network (CNN)” (Mandal, page 184, paragraph 3).
… wherein the persistence landscape for said each degree includes a collection of functions fi(x) = y from real numbers to real numbers: “For each of the simplicial complexes we obtained persistence landscapes [5] for homology dimensions 0 and 1. Such landscapes are, for each homology degree, sequences {
λ
k
} of decreasing piecewise linear (PL) functions
λ
k
:
R
→
R
” (Mandal, page 183, paragraph 3)
… wherein the tensor has a shape of I times X times Y times D, wherein I is a total number of functions fi, X is a total number of x values, Y is a total number of y values and D is a total number of the plurality of homology degrees considered: “In the TDA-CNN approach, for a given resolution
r
y
, we fed each subject’s vectorized persistence landscape as a tensor of shape (
r
x
,
r
y
, 2), one channel per homology degree, into a Convolutional Neural Network (CNN)” (Mandal, page 184, paragraph 3).
… the neural network being trained to predict a phenotype of a sample associated with a subject whose phenotype is unknown: “The topological summaries of the weighted point clouds (persistence landscapes) are then used to construct a machine learning model to predict the phenotype (healthy or PD) for each sample.” (Mandal, page 180, paragraph 3)
Mandal relates to using topological data analysis to diagnose diseased phenotypes with machine learning and is analogous to the claimed invention.
While Mandal fails to disclose the further limitations of the claim, Chazal discloses a method of subsampling data points of the gene expression data by randomly selecting a subset of a preconfigured size, the sub-sampling preserving a topology associated with the topological summaries:
“For any positive integer m, let
X
=
{
x
1
,
…
,
x
m
}
⊂
X
be a sample of m points from the measure
μ
∈
P
(
X
)
. The corresponding persistence landscape is
PNG
media_image1.png
321
879
media_image1.png
Greyscale
” Left: 3D shapes of the first experiment. Middle and Left: 500 random points from the magnetometer data of the second experiment.” (Chazal, page 7, Figure 3)
“In practice, each shape consists of a 3D point cloud embedded in the Euclidean space, with a number of vertices that ranges from 7K to 40K … For n = 100 times we subsample m = 300 points (preconfigured size) from each shape; then we select the closest subsample to the corresponding original point cloud and compute
4
×
n
persistence diagrams (dimension 1), one for each subsample” (Chazal, page 7, paragraph 3)
“We study the risk of two estimators and we prove that the subsampling approach carries stable topological (preserving a topology) information while achieving a great reduction in computational complexity” (Chazal, page 1, Abstract)
Mandal and Chazal relate to subsampling point clouds of data to reduce the cost of persistent homology analysis, and are analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Mandal to randomly subsample data such that subsamples maintain topological properties, as disclosed by Chazal. Chazal demonstrated that topological data analysis through these subsets accurately approximates persistent homology of the full set of data points, while being computationally faster and simple to execute. This is particularly useful when trying to perform topological data analysis is prohibitively expensive due to large data sets. See Chazal, page 1, Abstract & pages 7-9.
While Chazal fails to disclose the further limitations of the claim, Bubenik discloses a method, wherein the topological summaries including a persistence landscape for each degree of at least three homology degrees per subject of the plurality of subjects: “We sample 100 points from the uniform distribution on the unit cube [0, 1]^3, and calculate the persistence landscapes in degrees 0, 1 and 2 (homology degrees) of the corresponding Vietoris-Rips complex” (Bubenik, page 13, paragraph 4)
Bubenik relates to topological data analysis with machine learning and is analogous to the claimed invention. the existing combination teaches a method of training neural networks with persistence landscapes. Bubenik teaches a method of calculating persistence landscapes for homology degrees up to three. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the existing combination with Bubenik by using Bubenik’s method to calculate persistence landscapes. This would achieve the predictable result of calculating persistence landscapes for a finite number of homology degrees, with the existing combination’s method of training neural networks and Bubenik’s method of calculating persistence landscapes performing the same together as they did separately. (MPEP 2143 I. (A) Combining prior art elements according to known methods to yield predictable results).
While Bubenik fails to disclose the further limitations of the claim, Modarres discloses a method, wherein the tensor has a shape of I times X times Y times D, wherein I is a total number of functions fi: “The Net Table is a two-dimensional array (tensor) which contains a number of columns equal to the total number of functions through the hierarchy, and a number of rows equal to the number of nets specified by the user” (Modarres, column 11, paragraph 7)
Modarres is reasonably pertinent to the representation of functions in tensors and is analogous to the claimed invention. The existing combination teaches a method of representing persistence landscapes as three-dimensional tensors split along X values, Y values, and homology degrees. Modarres teaches tensors with a dimension for functions. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the existing combination and Modarres by additionally splitting persistence landscape tensors along a fourth dimension for landscape functions. This would achieve the predictable result of explicitly splitting the data along the persistence landscape function values, with the existing combination’s persistence landscape calculations and Modarres’ feature dimension performing the same together as they did separately. (MPEP 2143 I. (A) Combining prior art elements according to known methods to yield predictable results).
While Modarres fails to disclose the further limitations of the claim, Brandsma teaches [a] computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions readable by a device to: [execute operations]: “In embodiments, there is provided a non-transitory computer-readable medium having information recorded (instructions) thereon for generating a model for predicting severe disease in an individual with sepsis or at risk of developing sepsis, wherein the information, when read by a computer, causes the computer to perform operations” (Brandsma, [0020]).
Brandsma relates to neural networks and topological data analysis for disease phenotype prediction and is analogous to the claimed invention. The existing combination teaches a computational method of predicting disease phenotypes. The claimed invention improves upon this method by executing its method on computer hardware. Brandsma teaches computer hardware capable of running neural network methods for predicting phenotypes, applicable to the existing combination of the existing combination. A person of ordinary skill in the art would have recognized that running the existing combination’s method on Brandsma’s hardware would lead to the predictable result of the computational method being executed as described, and would improve the known device by allowing the method to produce concrete results on a computer (MPEP 2143 I. (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results).
The analysis of claims 17 and 19-20 mirrors that of claims 2 and 4-5, with the exception that claims 17 and 19-20 are directed to generic computer hardware which executes the methods of claims 2 and 4-5. This generic hardware is taught by Brandsma, as discussed regarding claim 16. Thus, claims 17 and 19-20 are rejected under the same rationales used for claims 2 and 4-5, respectively.
Claim 21 is rejected under 35 U.S.C. 103 as being unpatentable over Mandal (A Topological Data Analysis Approach on Predicting Phenotypes from Gene Expression Data, 2020, AlCoB 2020, LNBI 12099, pp. 178-187) in view of Chazal (Subsampling Methods for Persistent Homology, published 2014, arXiv:1406.1901v1), and further in view of Bubenik (STATISTICAL TOPOLOGICAL DATA ANALYSIS USING PERSISTENCE LANDSCAPES, published 1/23/2015, arXiv:1207.6437v4), Modarres (Hierarchical Floorplanner, published 4/17/1990, Patent no. 4,918,614), and Ferry (RECONSTRUCTING FUNCTIONS FROM RANDOM SAMPLES, 2014, Journal of Computational Dynamics, American Institute of Mathematical Sciences Volume 1, Number 2, December 2014).
Regarding claim 21, the rejection of claim 1 is incorporated. Mandal further discloses a method, wherein the transforming further includes resampling data of the gene expression data to replace the data points used in the topological summaries, wherein the data is resampled by enveloping each point by a sphere of a given radius and sampling points from a union of spheres, the resampling of the data preserving a topology associated with the topological summaries: “To mitigate the computational cost of our setup we used a subsampling approach, as studied in [11], so that instead of working with the entire set of genes at all times, for each subject we repeatedly subsampled smaller sets of
n
s
u
b
s
a
m
p
l
e
genes, obtaining several filtered simplicial complexes” (Mandal, page 183, paragraph 2). Subsampling is a type of resampling.
While the aforementioned references fail to disclose the further limitations of the claim, Ferry discloses a method, wherein the transforming further includes resampling data of the gene expression data to replace the data points used in the topological summaries, wherein the data is resampled by enveloping each point by a sphere of a given radius and sampling points from a union of spheres, the resampling of the data preserving a topology associated with the topological summaries:
“let U(X) denote the union of n-dimensional open ∈-balls (spheres) centered at the points in X” (Ferry, page 2, paragraph 3)
“Thus, the union of balls (union of spheres) of a suitably chosen radius around a sufficiently large point sample success to recover the homotopy type of that manifold with high confidence. From a computational perspective, recall that the nerve of a cover [12, 15] is the abstract simplicial complex where each d-dimensional simplex corresponds to an intersection of d + 1 sets of that cover. If we let N(X) denote the nerve corresponding to the cover of U(X) by its constituent open balls, then one obtains an isomorphism
H
*
(
χ
)
≃
H
*
∆
(
N
X
)
between the singular homology of 𝜒 and the simplicial homology of N(X)” (Ferry, page 2, paragraph 4). The homotopy type of the nerve and the original data manifold is the same, indicating they can be simply transformed into one another, preserving the topology of the original manifold. This is reinforced by the homology of the nerve and the original manifold being isomorphic.
“Nerves. Let U be a topological space equipped with a finite cover U (union of spheres) consisting of subsets (spheres) of M. The nerve of U is the simplicial complex N(U) with vertex set U where each subcollection
σ
⊂
U
constitutes a simplex if and only if the intersection
∩
u
∈
σ
is a non-empty subset of U” (Ferry, page 4, paragraph 3). The simplexes of the nerve are built only for intersections of spheres. In other words, intersections within the union of spheres are sampl[ed] to construct the simplicial complex of the nerve.
Ferry relates to topological reconstruction utilizing unions of spheres centered around data points and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the existing combination to transform the data into a simplicial complex, as disclosed by Ferry. Geometric properties are a useful way to provide insight to large high-dimensional datasets, but such data is typically finite and noisy. Topological invariants like homology and homotopy groups – two things measured by Ferry’s method – are useful ways to capture the underlying structure of this data. Additionally, Ferry’s method is resistant to noisy sample data. See Ferry, page 1, Abstract & page 1, paragraph 1.
Claims 22-23 are rejected under 35 U.S.C. 103 as being unpatentable over Mandal (A Topological Data Analysis Approach on Predicting Phenotypes from Gene Expression Data, 2020, AlCoB 2020, LNBI 12099, pp. 178-187) in view of Chazal (Subsampling Methods for Persistent Homology, published 2014, arXiv:1406.1901v1), and further in view of Bubenik (STATISTICAL TOPOLOGICAL DATA ANALYSIS USING PERSISTENCE LANDSCAPES, published 1/23/2015, arXiv:1207.6437v4), Modarres (Hierarchical Floorplanner, published 4/17/1990, Patent no. 4,918,614), Brandsma (Predicting and Addressing Severe Disease in Individuals with Sepsis, PCT filed 12/24/2020, US 2023/0018537 A1), and Ferry (RECONSTRUCTING FUNCTIONS FROM RANDOM SAMPLES, 2014, Journal of Computational Dynamics, American Institute of Mathematical Sciences Volume 1, Number 2, December 2014).
Regarding claim 22, the rejection of claim 9 is incorporated. Mandal further discloses a method, wherein the transforming further includes resampling data of the gene expression data to replace the data points used in the topological summaries, wherein the data is resampled by enveloping each point by a sphere of a given radius and sampling points from a union of spheres, the resampling of the data: “To mitigate the computational cost of our setup we used a subsampling approach, as studied in [11], so that instead of working with the entire set of genes at all times, for each subject we repeatedly subsampled smaller sets of 𝑛𝑠𝑢𝑏𝑠𝑎𝑚𝑝𝑙𝑒 genes, obtaining several filtered simplicial complexes” (Mandal, page 183, paragraph 2). Subsampling is a type of resampling.
While the aforementioned references fail to disclose the further limitations of the claim, Ferry discloses a method, wherein the transforming further includes resampling data of the gene expression data to replace the data points used in the topological summaries, wherein the data is resampled by enveloping each point by a sphere of a given radius and sampling points from a union of spheres, the resampling of the data:
“let U(X) denote the union of n-dimensional open ∈-balls (spheres) centered at the points in X” (Ferry, page 2, paragraph 3)
“Thus, the union of balls (union of spheres) of a suitably chosen radius around a sufficiently large point sample success to recover the homotopy type of that manifold with high confidence. From a computational perspective, recall that the nerve of a cover [12, 15] is the abstract simplicial complex where each d-dimensional simplex corresponds to an intersection of d + 1 sets of that cover. If we let N(X) denote the nerve corresponding to the cover of U(X) by its constituent open balls, then one obtains an isomorphism
H
*
(
χ
)
≃
H
*
∆
(
N
X
)
between the singular homology of 𝜒 and the simplicial homology of N(X)” (Ferry, page 2, paragraph 4). The homotopy type of the nerve and the original data manifold is the same, indicating they can be simply transformed into one another, preserving the topology of the original manifold. This is reinforced by the homology of the nerve and the original manifold being isomorphic.
“Nerves. Let U be a topological space equipped with a finite cover U (union of spheres) consisting of subsets (spheres) of M. The nerve of U is the simplicial complex N(U) with vertex set U where each subcollection
σ
⊂
U
constitutes a simplex if and only if the intersection
∩
u
∈
σ
is a non-empty subset of U” (Ferry, page 4, paragraph 3). The simplexes of the nerve are built only for intersections of spheres. In other words, intersections within the union of spheres are sampl[ed] to construct the simplicial complex of the nerve.
Ferry relates to topological reconstruction utilizing unions of spheres centered around data points and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the existing combination to transform the data into a simplicial complex, as disclosed by Ferry. Geometric properties are a useful way to provide insight to large high-dimensional datasets, but such data is typically finite and noisy. Topological invariants like homology and homotopy groups – two things measured by Ferry’s method – are useful ways to capture the underlying structure of this data. Additionally, Ferry’s method is resistant to noisy sample data. See Ferry, page 1, Abstract & page 1, paragraph 1.
Regarding claim 23, the rejection of claim 16 is incorporated. Mandal further discloses a method, wherein the transforming further includes resampling data of the gene expression data to replace the data points used in the topological summaries, wherein the data is resampled wherein the data is resampled by enveloping each point by a sphere of a given radius and sampling points from a union of spheres, the resampling of the data: “To mitigate the computational cost of our setup we used a subsampling approach, as studied in [11], so that instead of working with the entire set of genes at all times, for each subject we repeatedly subsampled smaller sets of 𝑛𝑠𝑢𝑏𝑠𝑎𝑚𝑝𝑙𝑒 genes, obtaining several filtered simplicial complexes” (Mandal, page 183, paragraph 2). Subsampling is a type of resampling.
While the aforementioned references fail to disclose the further limitations of the claim, Ferry discloses a method, wherein the transforming further includes resampling data of the gene expression data to replace the data points used in the topological summaries, wherein the data is resampled wherein the data is resampled by enveloping each point by a sphere of a given radius and sampling points from a union of spheres, the resampling of the data:
“let U(X) denote the union of n-dimensional open ∈-balls (spheres) centered at the points in X” (Ferry, page 2, paragraph 3)
“Thus, the union of balls (union of spheres) of a suitably chosen radius around a sufficiently large point sample success to recover the homotopy type of that manifold with high confidence. From a computational perspective, recall that the nerve of a cover [12, 15] is the abstract simplicial complex where each d-dimensional simplex corresponds to an intersection of d + 1 sets of that cover. If we let N(X) denote the nerve corresponding to the cover of U(X) by its constituent open balls, then one obtains an isomorphism
H
*
(
χ
)
≃
H
*
∆
(
N
X
)
between the singular homology of 𝜒 and the simplicial homology of N(X)” (Ferry, page 2, paragraph 4). The homotopy type of the nerve and the original data manifold is the same, indicating they can be simply transformed into one another, preserving the topology of the original manifold. This is reinforced by the homology of the nerve and the original manifold being isomorphic.
“Nerves. Let U be a topological space equipped with a finite cover U (union of spheres) consisting of subsets (spheres) of M. The nerve of U is the simplicial complex N(U) with vertex set U where each subcollection
σ
⊂
U
constitutes a simplex if and only if the intersection
∩
u
∈
σ
is a non-empty subset of U” (Ferry, page 4, paragraph 3). The simplexes of the nerve are built only for intersections of spheres. In other words, intersections within the union of spheres are sampl[ed] to construct the simplicial complex of the nerve.
Ferry relates to topological reconstruction utilizing unions of spheres centered around data points and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the existing combination to transform the data into a simplicial complex, as disclosed by Ferry. Geometric properties are a useful way to provide insight to large high-dimensional datasets, but such data is typically finite and noisy. Topological invariants like homology and homotopy groups – two things measured by Ferry’s method – are useful ways to capture the underlying structure of this data. Additionally, Ferry’s method is resistant to noisy sample data. See Ferry, page 1, Abstract & page 1, paragraph 1.
Response to Arguments
The following responses address arguments and remarks made in the instant remarks dated 06/15/2026.
103 Rejections
On page 9 of the instant remarks, the Applicant argues that the amended claims are not fully disclosed by the relied upon references:
“The cited references do not appear to disclose or suggest in particularity, the amended
features of claim 1. The same reasons apply to independent claims 9 and 16, and the pending
dependent claims at least by virtue of their dependencies. For those reasons, it is respectfully
requested that the rejection of the claims under this section be withdrawn.”
Regarding the Applicant’s arguments above, while the Examiner agrees that the previously relied upon references failed to fully disclose the amended claim limitations, upon further search and consideration, the claims are found to be obvious over the prior art in view of Mandal, Chazal, Bubenik, Modarres, Brandsma, Ferry.
In particular, regarding “wherein the persistence landscape for said each degree includes a collection of functions fi(x) = y from real numbers to real numbers”, Mandal discloses a collection of functions for each persistence landscape, each mapping real numbers to real numbers (Mandal, page 183, paragraph 3).
Regarding “wherein the tensor has a shape of I times X times Y times D, wherein I is a total number of functions fi, X is a total number of x values, Y is a total number of y values and D is a total number of the plurality of homology degrees considered”, Mandal discloses persistence landscape tensors of dimension X by Y by number of homology degrees (Mandal, page 184, paragraph 3). While Mandal fails to explicitly disclose forming tensors with a dimension for total number of functions, this deficiency is remedied by Modarres, which discloses forming tensors with a dimension for total number of functions (Modarres, column 11, paragraph 7).
Claims 1-5, 9-13, 16-17, and 19-23 are found to be obvious over the prior art and are rejected under 35 U.S.C. 103. See the 103 rejections section for more detail.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure
(Wu et al., published 2008, Network‐based global inference of human disease genes, Molecular Systems Biology (2008) 4: 189) teaches a method of using topological data analysis to build a predictive landscape for genotype-phenotype relationships for diseases
(Platt et al., published 2016, Characterizing redescriptions using persistent homology to isolate genetic pathways contributing to pathogenesis, Platt et al. BMC Systems Biology 2016, 10(Suppl 1):10) teaches a method of using persistent homology to identify genetic pathways for diseases
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Aaron P Gormley whose telephone number is (571)272-1372. The examiner can normally be reached Monday - Friday 12:00 PM - 8:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle T Bechtold can be reached on (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AG/Examiner, Art Unit 2148
/Ryan Barrett/Primary Examiner, Art Unit 2148