DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-4 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nava et al. (“Meta-Learning via Classifier(-free) Diffusion Guidance”, 31 Jan 2023) in view of Erkoc et al. (“HyperDiffusion: Generating Implicit Neural Fields with Weight-Space Diffusion”, 29 Mar 2023).
Regarding claim 1.
Nava teaches a computer-implemented method for deploying machine learning (ML) models for inference in production environments, the method being executed by one or more processors and comprising: see page 3, section 3, “for each training task Ti, we fine-tune the base model f(x,W) = y on it, and use the resulting Wi as a training sample to train a generative model of network weights”);
generating a latent representation of the set of weights by inputting the set of weights into an encoder of a variational autoencoder (VAE) (see page 4, “The fundamental building block of our unconditional generative model is the hypernetwork h(z,θ) = W that we can train in two ways: 1. HVAE: We define a Hypernet work VAE (Figure 2.A), which, given samples of fine-tuned base network weights Wi, learns a low-dimensional normally distributed latent representation zi. The encoder d(Wi,ω) = (µzi ,Σzi ) with parameters ω maps base network weights to means and variances used to sample a latent vector zi, while the decoder (or generator) is a classic hy per network h(zi,θ) = Wi which reconstructs the network weights from the latent vector (See Appendix A.6.1).”);
generating a denoised latent representation based on conditioning text by: diffusing the latent representation to generate a noisy latent representation, and denoising the noisy latent representation to provide the denoised latent representation (see page 2, figure 2a, “our Hypernetwork Latent Diffusion Model (HyperLDM) learns, conditional on the task embedding ei, to iteratively denoise a VAE latent vector zT i , . . . , z0 i over T iterations”, also see page 5, section 5.1, “Denoising Diffusion Probabilistic Models (Sohl-Dickstein et al., 2015; Ho et al., 2020, DDPM) are a powerful class of generative models designed to learn a data distribution p(x). They do so by learning the inverse of a forward diffusion process in which samples x0 of the data distribution are slowly corrupted with additive Gaussian noise over T steps with a variance schedule β1,...,βT…”, also see page 8, Generative Modeling and Classifier(-free) guidance, “Denoising Diffusion Probabilistic Models (Sohl-Dickstein et al., 2015; Ho et al., 2020, DDPM) overcome common issues in generative modeling using a simple likelihood-based reconstruction loss for iterative denoising, and have been shown to achieve state-of-the-art results in high resolution image generation (Dhariwal & Nichol, 2021; Rombach et al., 2022). Several techniques have been proposed for effective conditional sampling in generative and diffusion models, such as classifier/CLIP guidance” );
providing a reconstructed set of weights by inputting the denoised latent representation into a decoder of the VAE, the decoder outputting the reconstructed set of weights (see page 4, “1. HVAE: We define a Hypernet work VAE (Figure 2.A), which, given samples of fine-tuned base network weights Wi, learns a low-dimensional normally distributed latent representation zi. The encoder d(Wi,ω) = (µzi ,Σzi ) with parameters ω maps base network weights to means and variances used to sample a latent vector zi, while the decoder (or generator) is a classic hypernetwork h(zi,θ) = Wi which reconstructs the network weights from the latent vector (See Appendix A.6.1).”, also see page 17-18, Algorithm 8 HVAE Training, and Algorithm 13 HVAE + HyperLDM Training );
and deploying an updated ML model for production inference (see page 14, “The average performance of this base network with this fixed architecture when adapted and deployed on each individual test task is what allows us to fairly compare all meta-learning and multi-task algorithms. We acknowledge that, as a limitation of our work, the comparisons hold up when comparing relatively small fixed base networks f, and our approach might not be scalable to compete with massively pre-trained large scale multi-task models. In any case, we believe that weight space generation of compact models can be useful in a variety of contexts, such as when the adapted base model needs to be deployed in embedded systems and other domains with limited compute resources.”).
Nava does not specifically teach providing a ML model using a set of training data, the ML model having a set of weights associated therewith.
Erkoc teaches providing a ML model using a set of training data, the ML model having a set of weights associated therewith (see page 1, “Figure 1: HyperDiffusion enables a new paradigm in directly generating neural implicit fields by predicting their weight parameters. We leverage implicit neural fields to optimize a set of MLPs that faithfully represent individual dataset instances (“Overfitting,” top-left). Our network, based on a transformer architecture, then models a diffusion process directly on the optimized MLP weights (“Diffusion,” bottom-left). This enables synthesis of new implicit fields (“Synthesis,” bottom-right).”).
Both Nava and Erkoc pertain to the problem of optimizing neural network, thus being analogous. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine Nava and Erkoc to teach the above limitations. The motivation for doing so would be “we propose HyperDiffusion, a novel approach for unconditional generative modeling of implicit neural fields. HyperDiffusion operates directly on MLP weights and generates new neural implicit fields encoded by synthesized MLP parameters. Specifically, a collection of MLPs is first optimized to faithfully represent individual data samples. Subsequently, a diffusion process is trained in this MLP weight space to model the underlying distribution of neural implicit fields. HyperDiffusion enables diffusion modeling over a implicit, compact, and yet high-fidelity representation of complex signals across 3D shapes and 4D mesh animations within one single unified framework.” (see Erkoc abstract).
Regarding claim 2.
Nava and Erkoc teaches the method of claim 1,
Erkoc further teaches wherein the set of weights is provided as a weight matrix comprising weights that are generated for the ML model during training of the ML model (see page 4, “In the first MLP overfitting step, detailed in Section 4, we optimize a collection of MLPs such that each MLP rep resents a faithful neural occupancy field of a data sample (e.g., a 3D shape) from the training set. This enables highly accurate shape fitting due to the representation power of the neural fields. The optimized MLP weights are flattened into 1Dvectors and passed to a downstream diffusion process as ground truth signals.”).
The motivation utilized in the combination of claim 1, super, applies equally as well to claim 2.
Regarding claim 3.
Nava and Erkoc teaches the method of claim 1,
Nava further teaches wherein a dimension of the latent representation is less than a dimension of the set of weights (see page 4, “HVAE: We define a Hypernet work VAE (Figure 2.A), which, given samples of fine-tuned base network weights Wi, learns a low-dimensional normally distributed latent representation zi. The encoder d(Wi,ω) = (µzi ,Σzi ) with parameters ω maps base network weights to means and variances used to sample a latent vector zi, while the decoder (or generator) is a classic hypernetwork h(zi,θ) = Wi which reconstructs the network weights from the latent vector (See Appendix A.6.1).”).
Regarding claim 4.
Nava and Erkoc teaches the method of claim 1,
Nava further teaches wherein the conditioning text comprises a textual description of a condition that is one of absent from and underrepresented in the training data (see page 1, abstract, “We first train an unconditional generative hypernetwork model to produce neural network weights; then we train a second “guidance” model that, given a natural language task description, traverses the hypernetwork la tent space to find high-performance task-adapted weights in a zero-shot manner.”, also see page 3, “to perform zero-shot task adaptation, we utilize an additional high-level context embedding ei for each task Ti. In practice, such embeddings can come from a natural language description ti of the task, which can be encoded into a task embedding using pre-trained language models.”, also see page 4).
Claim(s) 5-7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Nava et al. (“Meta-Learning via Classifier(-free) Diffusion Guidance”, 31 Jan 2023) in view of Erkoc et al. (“HyperDiffusion: Generating Implicit Neural Fields with Weight-Space Diffusion”, 29 Mar 2023) in view of Malhotra et al. (“LSTM-based Encoder-Decoder for Multi-sensor Anomaly Detection”, Presented at ICML 2016 Anomaly Detection Workshop).
Regarding claim 5.
Nava and Erkoc teaches the method of claim 4,
Erkoc further teaches wherein the condition comprises one or more events see page 7, “In Figure 7, we visualize 4 out of 16 generated frames for both animation sequences. The full animation sequences in our video and website show that we can generate temporally consistent 4D animations corresponding to meaningful actions (e.g., jumping, rotating, and resting). Shape integrity is preserved throughout the generated animation sequence.”).
The motivation utilized in the combination of claim 1, super, applies equally as well to claim 5.
Nava and Erkoc do not teach one or more events associated with timeseries data.
Malhotra teaches one or more events associated with timeseries data (see page 2, section 2, “Consider a time-series X = {x(1),x(2),...,x(L)} of length L, where each point x(i) ∈ Rm is an m-dimensional vector of readings form variables at time-instance ti. We consider the scenario where multiple such time-series are available or can be obtained by taking a window of length L over a larger time-series. We first train the LSTM Encoder Decoder model to reconstruct the normal time-series. The reconstruction errors are then used to obtain the likelihood of a point in a test time-series being anomalous s.t. for each point x(i), an anomaly score a(i) of the point being anomalous is obtained. A higher anomaly score indicates a higher likelihood of the point being anomalous.”, also see page 4, “Space shuttle dataset contains periodic sequences with 1000 points per cycle, and 15 such cycles. We delibrately choose L = 1500 such that a subsequence covers more than one cycle (1.5 cycles per subsequence) and consider sliding windows with step size of 500. We down sample the original time-series by 3. The normal and anomalous sequences in Figure 3(c)-3(d) belong to TEK17 and TEK14 time-series, respectively.”, also see abstract, “We show that EncDec AD is robust and can detect anomalies from pre dictable, unpredictable, periodic, aperiodic, and quasi-periodic time-series. Further, we show that EncDec-AD is able to detect anomalies from short time-series (length as small as 30) as well as long time-series (length as large as 500).”, also see introduction, “The amount of load on a machine at a time may be unknown or change very frequently/abruptly, for example, in an earth digger. A machine may have multiple manual controls some of which may not be captured in the sensor data.”).
Nava, Erkoc and Malhotra pertain to the problem of optimizing neural network, thus being analogous. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine Nava, Erkoc and Malhotra to teach the above limitations. The motivation for doing so would be “We propose a Long Short Term Memory Networks based Encoder-Decoder scheme for Anomaly Detection (EncDec-AD) that learns to recon struct ‘normal’ time-series behavior, and there after uses reconstruction error to detect anomalies. We experiment with three publicly available quasi predictable time-series datasets: power demand, space shuttle, and ECG, and two real world engine datasets with both predictive and unpredictable behavior. We show that EncDec AD is robust and can detect anomalies from predictable, unpredictable, periodic, aperiodic, and quasi-periodic time-series. Further, we show that EncDec-AD is able to detect anomalies from short time-series (length as small as 30) as well as long time-series (length as large as 500).” (see Malhotra abstract).
Regarding claim 6.
Nava and Erkoc teaches the method of claim 1,
Nava further teaches wherein deploying the ML model comprises transmitting the ML model to a production environment to receive see page 14, “The average performance of this base network with this fixed architecture when adapted and deployed on each individual test task is what allows us to fairly compare all meta-learning and multi-task algorithms. We acknowledge that, as a limitation of our work, the comparisons hold up when comparing relatively small fixed base networks f, and our approach might not be scalable to compete with massively pre-trained large scale multi-task models. In any case, we believe that weight space generation of compact models can be useful in a variety of contexts, such as when the adapted base model needs to be deployed in embedded systems and other domains with limited compute resources.”).
Nava and Erkoc do not teach a production environment to receive timeseries data and generate inferences responsive to the timeseries data.
Malhotra teaches a production environment to receive timeseries data and generate inferences responsive to the timeseries data (see page 2, section 2, “Consider a time-series X = {x(1),x(2),...,x(L)} of length L, where each point x(i) ∈ Rm is an m-dimensional vector of readings form variables at time-instance ti. We consider the scenario where multiple such time-series are available or can be obtained by taking a window of length L over a larger time-series. We first train the LSTM Encoder Decoder model to reconstruct the normal time-series. The reconstruction errors are then used to obtain the likelihood of a point in a test time-series being anomalous s.t. for each point x(i), an anomaly score a(i) of the point being anomalous is obtained. A higher anomaly score indicates a higher likelihood of the point being anomalous.”, also see page 4, “Space shuttle dataset contains periodic sequences with 1000 points per cycle, and 15 such cycles. We delibrately choose L = 1500 such that a subsequence covers more than one cycle (1.5 cycles per subsequence) and consider sliding windows with step size of 500. We down sample the original time-series by 3. The normal and anomalous sequences in Figure 3(c)-3(d) belong to TEK17 and TEK14 time-series, respectively.”, also see abstract, “We show that EncDec AD is robust and can detect anomalies from pre dictable, unpredictable, periodic, aperiodic, and quasi-periodic time-series. Further, we show that EncDec-AD is able to detect anomalies from short time-series (length as small as 30) as well as long time-series (length as large as 500).”, also see introduction, “The amount of load on a machine at a time may be unknown or change very frequently/abruptly, for example, in an earth digger. A machine may have multiple manual controls some of which may not be captured in the sensor data.”).
Nava, Erkoc and Malhotra pertain to the problem of optimizing neural network, thus being analogous. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine Nava, Erkoc and Malhotra to teach the above limitations. The motivation for doing so would be “We propose a Long Short Term Memory Networks based Encoder-Decoder scheme for Anomaly Detection (EncDec-AD) that learns to recon struct ‘normal’ time-series behavior, and there after uses reconstruction error to detect anomalies. We experiment with three publicly available quasi predictable time-series datasets: power demand, space shuttle, and ECG, and two real world engine datasets with both predictive and unpredictable behavior. We show that EncDec AD is robust and can detect anomalies from predictable, unpredictable, periodic, aperiodic, and quasi-periodic time-series. Further, we show that EncDec-AD is able to detect anomalies from short time-series (length as small as 30) as well as long time-series (length as large as 500).” (see Malhotra abstract).
Regarding claim 7.
Nava and Erkoc teaches the method of claim 1,
Nava and Erkoc do not teach claim 7,
Malhotra teaches wherein the ML model is a long short-term memory autoencoder (LSTM-AE) (see abstract, “We propose a Long Short Term Memory Networks based Encoder-Decoder scheme for Anomaly Detection (EncDec-AD) that learns to recon struct ‘normal’ time-series behavior, and there after uses reconstruction error to detect anomalies…We show that EncDec AD is robust and can detect anomalies from pre dictable, unpredictable, periodic, aperiodic, and quasi-periodic time-series. Further, we show that EncDec-AD is able to detect anomalies from short time-series (length as small as 30) as well as long time-series (length as large as 500).”, also see introduction, “An LSTM-based en coder is used to map an input sequence to a vector representation of fixed dimensionality. The decoder is another LSTM network which uses this vector representation to produce the target sequence.”, also see page 2, section 2.1).
Nava, Erkoc and Malhotra pertain to the problem of optimizing neural network, thus being analogous. It would have been obvious to one skilled in the art before the effective filing date of the claimed invention to combine Nava, Erkoc and Malhotra to teach the above limitations. The motivation for doing so would be “We propose a Long Short Term Memory Networks based Encoder-Decoder scheme for Anomaly Detection (EncDec-AD) that learns to recon struct ‘normal’ time-series behavior, and there after uses reconstruction error to detect anomalies. We experiment with three publicly available quasi predictable time-series datasets: power demand, space shuttle, and ECG, and two real world engine datasets with both predictive and unpredictable behavior. We show that EncDec AD is robust and can detect anomalies from predictable, unpredictable, periodic, aperiodic, and quasi-periodic time-series. Further, we show that EncDec-AD is able to detect anomalies from short time-series (length as small as 30) as well as long time-series (length as large as 500).” (see Malhotra abstract).
Claims 8-14 recites a non-transitory computer-readable storage medium to perform the method recited in claims 1-7. Therefore, the rejection of claims 1-7 above applies equally here.
Claims 15-20 recites a system to perform the method recited in claims 1-7. Therefore, the rejection of claims 1-7 above applies equally here.
Related prior arts:
Luo et al. (“Understanding Diffusion Models: A Unified Perspective”, Google Research, Brain Team, 2022) teaches a distribution is learned as an arbitrarily flexible energy function that is then normalized. 1 Score-based generative models are highly related; instead of learning to model the energy function itself, they learn the score of the energy-based model as a neural network. In this work we explore and review diffusion models, which as we will demonstrate, have both likelihood-based and score-based interpretations. We showcase the math behind such models in excruciating detail, with the aim that anyone can follow along and understand what diffusion models are and how they work.
Hafner et al. (US 20210158162 A1) teaches receiving a latent representation characterizing a current state of the environment; generating a trajectory of latent representations that starts with the received latent representation; for each latent representation in the trajectory: determining a predicted reward; and processing the state latent representation using a value neural network to generate a predicted state value; determining a corresponding target state value for each latent representation in the trajectory; determining, based on the target state values, an update to the current values of the policy neural network parameters; and determining an update to the current values of the value neural network parameters.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to IMAD M KASSIM whose telephone number is (571)272-2958. The examiner can normally be reached 10:30AM-5:30PM, M-F (E.S.T.).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michael J. Huntley can be reached at (303) 297 - 4307. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/IMAD KASSIM/Primary Examiner, Art Unit 2129