Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 05/02/2024 was filed before the mailing date of the first office action. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 4 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claim 4 recites the limitation “The device according to any one of the preceding claims, wherein n is equal to 1 and said input n-tuples of features are single real numbers”. The dependency of claim 4 is unclear and consequently it is also unclear what the variable “n” is meant to refer to. For purposes of examination, Examiner is interpreting that claim 4 depends on claim 1 and that the variable n refers to the “input n-tuples” from claim 1 as opposed to the “n instances of said base model” from claim 2.
The following is a quotation of 35 U.S.C. 112(d):
(d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph:
Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers.
Claim 13 is rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. All the limitations in claim 13 are recited in claim 12; therefore, claim 13 cannot be said to further limit claim 1 on which both claims 12 and 13 depend. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-16 are rejected under 35 U.S.C. 103 as being unpatentable over Lesh et al* (US 20230010686 A1, herein Lesh) in view of Canale et al (“Generative Modeling of Complex Data”, herein Canale).
*this document was included in the IDS dated 05/02/2024
Regarding claim 1, Lesh teaches a device for training a privacy-preserving generative model for management of data privacy configured to generate a synthetic time series, at least one processor configured to train an untrained generative model based on said training dataset and by using a privacy-preserving training technique, so as to obtain model parameters for said trained privacy-preserving generative model configured to output at least one synthetic time series, at least one output configured to output said model parameters associated to said trained privacy-preserving generative model (para. [0004] recites “Systems, methods, and articles of manufacture, including computer program products, are provided for preparing data for machine learning processing and synthetic data generation”. Para. [0040] recites “Since multiple events may be created per synthetic patient, the synthetic data may be longitudinal, that is, not a set of static characteristics of a patient such as age, gender, diagnoses, but a complete patient trajectory of medical events over time”. Para. [0074] recites “Some implementations may utilize a federated learning architecture” (i.e., using a federated, or privacy-preserving, generative model to generate synthetic time-series data)),
wherein: said generative model comprises a combination of: a first individual generative model configured to receive as input said input length values, a second individual generative model configured to receive as input said input timestamps (para. [0059] recites “electronic health record data may be in any format and may include patient medical events. Electronic health record data may be viewed as a sequence of medical events. In each medical event, a timestamp along with one or more medically relevant information segments may be recorded about a patient. Electronic health record data may be structured or unstructured, may be numeric, textual, image, video, or the like, may contain continuous or categorical or binary values, as well as missing values (null), may be stationary or time varying, or may be a single data item or sequence of data values over time”. Para. [0072] recites “The transformer architecture's encoder may include a stack of N encoders (e.g., typically N = 6 is used), and similarly the transformer architecture may include a stack of N decoders/generators” (i.e., a combination of generative models can be used to process at least continuous, or input length, values and timestamp data)).
However, Lesh does not explicitly teach a third individual generative model configured to receive as input said input data of the structure type, a first causal transformer block configured to receive as input a plurality of embedding vectors obtained from said first, second and third individual generative models, a second causal transformer block configured to receive as input a conditioning vector associated to said generative model, and a plurality of compressed representations obtained from said first causal transformer block.
Canale teaches a third individual generative model configured to receive as input said input data of the structure type (section I recites “We introduce a new framework that systematically maps a large class of data types to generative models called codecs. We explicit the codecs architecture for primitive types (categorical and numerical data), and composite (structs and lists) and show that in particular these codecs allow us to synthesize standard and complex hierarchical tabular datasets”. Section 3.3.1 recites “An observation x of struct type is a tuple of features (x1, x2, . . ., xn), associated with labels” (i.e., a generative model for structure type input data)),
a first causal transformer block configured to receive as input a plurality of embedding vectors obtained from said first, second and third individual generative models (section 3.3.1 recites “A causal transformer block (see A.3) is then applied to the embedding vectors, yielding a list of representations used to build the intermediate context and the struct embedding vector” (i.e., a causal transformer block can be applied to the embedding vectors in each generative model)),
a second causal transformer block configured to receive as input a conditioning vector associated to said generative model, and a plurality of compressed representations obtained from said first causal transformer block (Canale section 3.3.1 recites “The conditioning vector is combined with part of the intermediate context using another causal transformer”. Canale section A.3 recites “the transformer was initially made of multiple blocks based on the GPT2 architecture. Since the transformer is not very deep (typically 2 blocks are enough), we found experimentally that removing the layer normalization and the dense did not alter the results” (i.e., a second causal transformer block is used)).
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine these teachings by utilizing the privacy preserving generative model framework from Canale to implement the synthetic time-series data generation from Lesh. Lesh and Canale are both directed to generative models, and one of ordinary skill in the art would be motivated to generate the time-series data from Lesh using the model from Canale, as Canale teaches in at least section 2 that tabular time-series data can be modeled using transformer frameworks.
Lesh in view of Canale teaches said synthetic time series being defined by a length, a sequence of timestamps (para. [0059] recites “electronic health record data may be in any format and may include patient medical events. Electronic health record data may be viewed as a sequence of medical events. In each medical event, a timestamp along with one or more medically relevant information segments may be recorded about a patient. Electronic health record data may be structured or unstructured, may be numeric, textual, image, video, or the like, may contain continuous or categorical or binary values, as well as missing values (null), may be stationary or time varying, or may be a single data item or sequence of data values over time” (i.e. length and timestamp time series data)), and sequence of data of structure type, each data of structure type comprising a n-tuple of features, associated to one timestamp in the sequence of timestamps (Canale section 3.3.1 recites “An observation x of struct type is a tuple of features (x1, x2, . . ., xn), associated with labels” (i.e., structure data comprising an n-tuple of features, which could be associated to the timestamped event data from Lesh)), said device comprising:
at least one input configured to receive a training dataset comprising a set of private input time series, each private input time series among said set of private input time series comprising an input length value mk, a sequence of mk input timestamps (Lesh para. [0059] recites “electronic health record data may be in any format and may include patient medical events. Electronic health record data may be viewed as a sequence of medical events. In each medical event, a timestamp along with one or more medically relevant information segments may be recorded about a patient. Electronic health record data may be structured or unstructured, may be numeric, textual, image, video, or the like, may contain continuous or categorical or binary values, as well as missing values (null), may be stationary or time varying, or may be a single data item or sequence of data values over time”. Lesh para. [0074] recites “Some implementations may utilize a federated learning architecture” (i.e. the federated, or privacy-preserving model, can process length and timestamp time series input data)), a sequence of mk input data of structure type, each input data comprising an n-tuple of features, associated to one input timestamp among the corresponding sequence of mk input timestamps (Canale section 3.3.1 recites “An observation x of struct type is a tuple of features (x1, x2, . . ., xn), associated with labels” (i.e., the federated, or privacy-preserving model from Lesh can process data like the structure data comprising an n-tuple of features from Canale, which could be associated to the timestamped event data from Lesh));
said trained privacy-preserving generative model is configured to model relationships between said input length values, said input timestamps and said input data of the structure type (Canale section 4.1 recites “To evaluate feature correlations, we compute the norm of the difference between the pair-wise correlation matrices (real and synthetic)” (i.e., the federated, or privacy-preserving model from Lesh can model correlations, or relationships, between input data such as the length and timestamp data from para. [0059] and the structure type data from Canale)).
Regarding claim 2, the combination of Lesh and Canale teaches the device of claim 1 as mentioned above, wherein each among said first individual generative model and said second individual generative model comprises an instance of a base model, and said third individual generative model comprises a combination of n instances of said base model (Canale fig. 2 and section 3.3.1 recite “We can define a struct codec Cstruct [C1,C2, . . . ,Cn] by combining n feature codecs Ck ϵ (Ek, Dk, Sk, Lk), as illustrated in Fig. 2-A and Fig. 6 of A.1”. Lesh para. [0062] recites “Transforming the electronic health record into a numerical vector may include mapping the event-code to a first vector (of N dimension), normalizing and embedding the event-value (if one exists) into a second vector (of M dimension), and concatenating the first vector and the second vector into a final N+M dimensional vector” (i.e., multiple generative models can be used together in sequence or in combination)),
wherein said base model is defined by a combination of an encoder, a decoder, a loss function, a sampler and a conditioning vector (Canale section 3.1.1 recites “A codec is a generative model defined by a quadruplet: C = (E, D, S, L) and a fixed initial conditioning vector c0 ϵ E ” (i.e., a model with a combination of E/encoder, D/decoder, S/sampler, L/loss function, and a conditioning vector)),
wherein: said encoder is configured to encode an input element into an embedding vector and a compressed representation (Canale section 3.1.1 recites “The encoder takes an observation x and returns a pair of an embedding vector summarizing everything there is to know about x, and an intermediate context containing all intermediate information about a composite observation while being processed” (i.e., an encoder to encode an input into an embedding vector and an intermediate context, or compressed representation)),
said decoder is configured to receive said conditioning vector and said compressed representation and to output a distribution representation (Canale section 3.1.1 recites “The decoder takes a conditioning vector from the embedding space and an intermediate context, and returns a distribution representation” (i.e., a decoder to decode the conditioning vector and intermediate representation, or compressed representation to output a distribution representation)),
said sampler is configured to receive said conditioning vector and said distribution representation and to output an output element (Canale section 3.1.1 recites “Given some probability space, the sampler is a random variable taking a conditioning vector, an outcome, and returning a sampled observation” (i.e., a sampler to receive the conditioning vector and an outcome, or distribution representation from the decoder)),
said loss function is defined based on said distribution representation, said input element and said output element (Canale section 3.1.1 recites “The loss function is a measurable function taking a decoded distribution representation, an actual observation, an outcome and returning a loss” (i.e., a loss function defined by the distribution representation from the decoder, the input observation, and an output from the sampler)).
Regarding claim 3, the combination of Lesh and Canale teaches the device of claim 1 as mentioned above, wherein said input timestamps are unevenly distributed (Lesh para. [0069] recites “it may be the case that a single patient's medical event distribution in time is highly non-uniform as the gap between events can be in hours, days or even years” (i.e., timestamps can be uneven)).
Regarding claim 4, the combination of Lesh and Canale teaches the device of any one of the preceding claims as mentioned above, wherein n is equal to 1 and said input n-tuples of features are single real numbers (Examiner notes that this claim is interpreted such that it depends on claim 1 and that the variable n refers to the input n-tuples from claim 1. Given this interpretation, at least section 3.2.1 of Canale teaches a categorical codec model for “An observation x of categorical type is an element of a finite set X = {a1, a2, . . ., an} of cardinality n”).
Regarding claim 5, the combination of Lesh and Canale teaches the device of claim 2 as mentioned above, wherein said at least one processor is configured to train said untrained generative model using said privacy-preserving training technique by: encoding, in a parallel manner, for a subset of private input time series among said set of private input time series, the corresponding input lengths by said encoder of said first individual generative model, the corresponding input timestamps by said encoder of said second individual generative model and the corresponding input data by said n encoders of said third individual generative model in a sequential manner, so as to obtain a corresponding subset of embedding vectors and a corresponding subset of compressed representations referred to as overall compressed representation (Canale fig. 2 and section 3.1.1 recites “The encoder takes an observation x and returns a pair of an embedding vector summarizing everything there is to know about x, and an intermediate context containing all intermediate information about a composite observation while being processed”. Lesh para. [0072] recites “The transformer architecture's encoder may include a stack of N encoders (e.g., typically N = 6 is used), and similarly the transformer architecture may include a stack of N decoders/generators” (i.e., the generative models can encode data in parallel such as the time-series length and time-stamped data from Lesh)),
outputting a subset of distribution representations, in a parallel manner, using said decoders of said first, second and third individual generative model and based on said subset of embedding vectors and on a corresponding augmented subset of compressed representations, obtained from part of said overall compressed representation (Canale fig. 2 and section 3.1.1 recites “The decoder takes a conditioning vector from the embedding space and an intermediate context, and returns a distribution representation” (i.e., the generative models can output distribution representations based on the embedding vectors and compressed representations from the encoders)),
minimizing an overall loss function based the loss functions corresponding to respectively the first, the second and the third individual generative models by respectively modifying said first individual generative model, said second individual generative model and said third individual generative model, said first causal transformer block and said second causal transformer block (Canale section 3.1.1 recites “A codec is a generative model defined by a quadruplet: C = (E, D, S, L) and a fixed initial conditioning vector c0 ϵ E ”. The loss function is a measurable function taking a decoded distribution representation, an actual observation, an outcome and returning a loss” (i.e., a loss function defined by the distribution representation from the decoder, the input observation, and an output from the sampler)),
repeating encoding, outputting and minimizing for another subset of private input time series among said set of private input time series until encoding has been applied to all private input time series of said set of private input time series (Lesh para. [0067] recites “the generator 605, the discriminator 610, and/or the encoder 615 may undergo multiple of training iterations” (i.e., training steps can be iterated, or repeated)).
Regarding claim 6, the combination of Lesh and Canale teaches the device of claim 5 as mentioned above, wherein said at least one processor is configured to train said untrained generative model by carrying out of minimizing said overall loss function via a differentially private stochastic gradient descent algorithm (Canale appendix C recites “we provide some preliminary work to show the promises of our model trained with differential privacy. In fact, the model can be easily trained using the standard algorithm of DP-SGD” (i.e., the model can be trained used differentially private stochastic gradient descent)).
Regarding claim 7, the combination of Lesh and Canale teaches the device of claim 5 as mentioned above, wherein said augmented set of compressed representations is obtained based on the application of said first causal transformer block to said set of embedding vectors (Canale section 3.3.1 recites “A causal transformer block (see A.3) is then applied to the embedding vectors, yielding a list of representations used to build the intermediate context and the struct embedding vector” (i.e., a causal transformer block can be applied to the embedding vectors in each generative model)).
Regarding claim 8, the combination of Lesh and Canale teaches the device of claim 7 as mentioned above, wherein said augmented set of compressed representations is further obtained based on the application of said second causal transformer block to the output of the application of said first causal transformer block to said set of embedding vectors (Canale section 3.3.1 recites “The conditioning vector is combined with part of the intermediate context using another causal transformer”. Canale section A.3 recites “the transformer was initially made of multiple blocks based on the GPT2 architecture. Since the transformer is not very deep (typically 2 blocks are enough), we found experimentally that removing the layer normalization and the dense did not alter the results” (i.e., a second causal transformer block is used)).
Regarding claim 9, the combination of Lesh and Canale teaches the device of claim 1 as mentioned above, wherein said input data in said sequence of input data comprise one among storage management data, fleet management data, personal activity tracking data, autonomous vehicles data or medical records (Lesh para. [0032] recites “The system and methods described herein may be used to generate a synthetic electronic health record dataset or augment an existing electronic health record dataset to make it more usable for downstream applications”).
Claim 10 is a method claim and its limitation is included in claim 1. The only difference is that claim 10 requires a method (Lesh para. [0004] recites “Systems, methods, and articles of manufacture, including computer program products, are provided for preparing data for machine learning processing and synthetic data generation”). Therefore, claim 10 is rejected for the same reasons as claim 1.
Regarding claim 11, the combination of Lesh and Canale teaches a non-transitory program storage device, readable by a computer, tangibly embodying a program of instructions executable by the computer to perform a method for training according to claim 1 (Lesh para. [0006] recites “A memory, which can include a non-transitory computer-readable or machine-readable storage medium, may include, encode, store, or the like one or more programs that cause one or more processors to perform one or more of the operations described herein”).
Regarding claim 12, the combination of Lesh and Canale teaches a device for management of data privacy configured to generate synthetic time series using a trained privacy-preserving generative model parametrized with model parameters obtained by a device for training according to claim 1 (Lesh para. [0004] recites “Systems, methods, and articles of manufacture, including computer program products, are provided for preparing data for machine learning processing and synthetic data generation”. Para. [0040] recites “Since multiple events may be created per synthetic patient, the synthetic data may be longitudinal, that is, not a set of static characteristics of a patient such as age, gender, diagnoses, but a complete patient trajectory of medical events over time”. Para. [0074] recites “Some implementations may utilize a federated learning architecture”. Canale section 3.1.1 recites “A codec is usually parametrized so that it can be fitted to real data” (i.e., the federated, or privacy-preserving model from Lesh can be trained to generate synthetic time-series data based model parameters such as the parameterized data from Canale)),
said synthetic time series being defined by a length, a sequence of timestamps and a sequence of data of structure type, each data of structure type comprising a n-tuple of features, associated to one timestamp in the sequence of timestamps, said device comprising: at least one input configured to receive said model parameters (Lesh para. [0059] recites “electronic health record data may be in any format and may include patient medical events. Electronic health record data may be viewed as a sequence of medical events. In each medical event, a timestamp along with one or more medically relevant information segments may be recorded about a patient. Electronic health record data may be structured or unstructured, may be numeric, textual, image, video, or the like, may contain continuous or categorical or binary values, as well as missing values (null), may be stationary or time varying, or may be a single data item or sequence of data values over time” (i.e. length and timestamp time series data). Canale section 3.3.1 recites “An observation x of struct type is a tuple of features (x1, x2, . . ., xn), associated with labels” (i.e., structure data comprising an n-tuple of features, which could be associated to the timestamped event data from Lesh)),
at least one processor configured to: generate said synthetic time series using said trained privacy-preserving generative model parametrized with said model parameters, at least one output configured to output said synthetic time series (Lesh para. [0032] recites “The system and methods described herein may be used to generate a synthetic electronic health record dataset or augment an existing electronic health record dataset to make it more usable for downstream applications” (i.e., the model can output, or generate synthetic time-series data using the federated, or privacy-preserving model from para. [0074] of Lesh)).
As noted in the 112(d) rejection, all the limitations of claim 13 are recited in claim 12. Claim 13 is rejected for the same reasons as claim 12.
Regarding claim 14, the combination of Lesh and Canale teaches the device according to claim 11 as mentioned above, wherein said at least one processor is configured to sample said synthetic temporal series in an autoregressive manner (Canale section 3.3.1 recites “The sampler Slist takes a conditioning context, samples the number m of repetitions then sequentially samples m values. The values are sampled by encoding the length and each already sampled feature and drawing autoregressively”. Lesh para. [0040] recites “the synthetic data may be longitudinal, that is, not a set of static characteristics of a patient such as age, gender, diagnoses, but a complete patient trajectory of medical events over time” (i.e., the data, such as the synthetic time series data from Lesh, can be sampled autoregressively)).
Claim 15 is a method claim and its limitation is included in claim 12. The only difference is that claim 15 requires a method (Lesh para. [0004] recites “Systems, methods, and articles of manufacture, including computer program products, are provided for preparing data for machine learning processing and synthetic data generation”). Therefore, claim 15 is rejected for the same reasons as claim 12.
Regarding claim 16, the combination of Lesh and Canale teaches a non-transitory program storage device, readable by a computer, tangibly embodying a program of instructions executable by the computer to perform a method of data privacy according to claim 15 (Lesh para. [0006] recites “A memory, which can include a non-transitory computer-readable or machine-readable storage medium, may include, encode, store, or the like one or more programs that cause one or more processors to perform one or more of the operations described herein”).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
“Composable Generative Models” (Leduc et al) teaches a method for conditional generative modeling of tabular data with privacy preserving applications.
“Real-valued (Medical) Time Series Generation with Recurrent Conditional GANs” (Esteban et al) teaches a Recurrent GAN (RGAN) and Recurrent Conditional GAN (RCGAN) to produce realistic real-valued multi-dimensional time series, with an emphasis on their application to medical data.
“Stochastic gradient descent with differentially private updates” (Song et al) teaches a derivation of differentially private stochastic gradient descent.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to LEAH M FEITL whose telephone number is (571) 272-8350. The examiner can normally be reached on M-F 0900-1700 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached on (571) 270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/L.M.F./ Examiner, Art Unit 2147
/VIKER A LAMARDO/Supervisory Patent Examiner, Art Unit 2147