Generative model handles varying data dimensions using jump diffusion processes.
problem Handling data of varying dimensionality in generative models.
method Formulated as a jump diffusion process, learning to approximate the process with a novel evidence lower bound.
result Effective sampling of data of varying dimensionality, better compatibility with test-time diffusion guidance imputation tasks.
Researchers formalize PD and PFI to relate them to data generating process.
problem Lack of theory linking PD and PFI to data generating process.
method Formalize PD and PFI as estimators of ground truth estimands, account for model variance with learner-PD and learner-PFI.
result PD and PFI estimates deviate from ground truth due to statistical biases, model variance, and Monte Carlo approximation errors.
Generative model on manifolds reduces divergence computation and improves scalability.
problem Difficulties in modeling data on non-Euclidean spaces due to expensive divergence computation and approximations of heat kernel.
method Riemannian Diffusion Mixture, a principled framework using a mixture of bridge processes.
result Achieves superior performance on diverse manifolds with reduced simulation steps.
A method for generating future samples in auto-regressive processes using confidence-based sampling.
problem Generating future samples in auto-regressive processes efficiently and accurately.
method Confidence-based sampling to simplify and parallelize the auto-regressive generation process.
result The method successfully captures complex data structures and generates meaningful future samples with lower computational cost.
Paper reviews multi-way graph signal processing for tensor data.
problem Maximizing use of multi-way structure in irregular tensor data.
method Generalizes GSP to multi-way data, focusing on graph signals across tensor modes.
result Synthesizes common themes in combining GSP with tensor analysis.
New methods improve estimation of nonhomogeneous Poisson processes from limited data.
problem Estimating nonhomogeneous Poisson processes from limited data.
method Formulated as a learning generalization problem, proposed adaptive and data-driven binning methods.
result Improved estimation of nonhomogeneous Poisson processes with limited data.
Paper benchmarks machine learning for detecting process curve drifts.
problem Detecting drifts in multivariate manufacturing process data.
method Synthetic data generation and evaluation score introduction.
result Existing algorithms often fail with complex drift scenarios.
GP-ConvCNP improves NP models for time series data by adding Gaussian Process.
problem GP-ConvCNP addresses the lack of generalization and robustness in ConvCNP models for time series data.
method GP-ConvCNP incorporates a Gaussian Process to improve ConvCNP's performance and generalization.
result GP-ConvCNP models show improved generalization and robustness to distribution shifts and future extrapolation.
Paper introduces a neural network-based non-stationary influence kernel for complex event data.
problem Modeling complex, non-stationary, and dependent discrete event data.
method Neural Spectral Marked Point Processes (NSMPP) with a versatile non-stationary influence kernel.
result NSMPP outperforms state-of-the-art models on synthetic and real data.
Proposes LDIDPs for efficient sequential data generation from latent dynamical models.
problem Challenges in generating high-fidelity sequential samples from latent dynamical models.
method Utilizes implicit diffusion processes to sample from latent dynamical processes.
result Demonstrates accurate learning of dynamics and efficient generation of high-quality sequential data.
Framework evaluates privacy cost of non-private pre-processing in DP pipelines.
problem Privacy cost of non-private data-dependent pre-processing in DP machine learning pipelines.
method Establishes upper bounds on overall privacy guarantees using Smooth DP and bounded sensitivity.
result Explicit overall privacy guarantees for various pre-processing algorithms.
Process mining is a research field focused on the analysis of event data with the aim of extracting insights in processes. Applying process mining techniques on data from smart home environments has the potential to provide valuable insights in (un)healthy habits and to contribute to ambient assisted living solutions. …
RML improves generative modeling of complex distributions.
problem Learning complex distributions in applications.
method RML defines a forward process to a known distribution, then learns a reverse Markov process.
result RML efficiently captures complex distributions in simulations and climate data.
New test for point processes without strong model assumptions.
problem Testing local independence in point processes without strong model assumptions.
method Expansion similar to Volterra expansions to represent marginalized intensities.
result Approximation of true marginalized intensity arbitrarily well.
This article outlines a method for automatically generating models of dynamic decision-making that both have strong predictive power and are interpretable in human terms. This is useful for designing empirically grounded agent-based simulations and for gaining direct insight into observed dynamic processes. We use an e…
TSFlow uses Gaussian processes to match priors for better time series forecasting.
problem Difficulties in aligning generative models' priors with time series data.
method Conditional flow matching (CFM) with Gaussian processes, optimal transport, and data-dependent priors.
result TSFlow produces high-quality unconditional samples and competitive forecasting results.
The data association problem is concerned with separating data coming from different generating processes, for example when data come from different data sources, contain significant noise, or exhibit multimodality. We present a fully Bayesian approach to this problem. Our model is capable of simultaneously solving the…
A novel model learns from limited data using physics constraints and GPVAE to generate realistic samples.
problem Limited data for effective generative AI training.
method Physics-informed Gaussian Process Variational Autoencoder (PIGPVAE) incorporating physical models and discrepancy terms.
result Achieves state-of-the-art performance on indoor temperature data.
The study addresses overlooked data-generating processes in time-series asset pricing.
problem The literature on time-series asset pricing overlooks the data-generating processes for factors expressed in return differences.
method The study proposes a new definition of returns and compound returns for factors, and uses OLS with net returns for single-index models.
result OLS with net returns for single-index models leads to inflated alphas, exaggerated t-values, and overestimated Sharpe ratios.
A new method improves the interpretability of data-driven models in ironmaking processes.
problem Lack of transparency in machine learning models used in industrial processes.
method Combines Variational Autoencoder (VAE) with Local Interpretable Model-agnostic Explanations (LIME) for better model interpretability.
result Improved local fidelity of local interpretable linear models compared to LIME.
Computer simulations have become a popular tool of assessing complex skills such as problem-solving skills. Log files of computer-based items record the entire human-computer interactive processes for each respondent. The response processes are very diverse, noisy, and of nonstandard formats. Few generic methods have b…
Paper introduces DMPMs for efficient discrete data generation with sharp convergence bounds.
problem Efficient generation of discrete data with theoretical guarantees.
method Discrete Markov Probabilistic Models (DMPMs) operating in bit space with time-reversal process.
result Sharp convergence bounds established under minimal assumptions, competitive performance in discrete data generation.
This work extends diffusion models to function space for better generative modeling.
problem Limited applicability of diffusion models to functional data domains.
method Introduces Denoising Diffusion Operators (DDOs) for training diffusion models in function space.
result Demonstrates accurate function-valued generation at fixed cost.
Scientific and engineering processes deliver massive high-dimensional data sets that are generated as non-linear transformations of an initial state and few process parameters. Mapping such data to a low-dimensional manifold facilitates better understanding of the underlying processes, and enables their optimization. I…
Develops a new method for functional regression that works with non-Gaussian data.
problem Limited models for regression in function spaces with Gaussian process priors.
method Introduces Neural Operator Flows (OpFlow) for non-Gaussian function spaces.
result OpFlow enables robust and accurate uncertainty quantification for functional regression.
Proposes first privacy-preserving method for estimating Hawkes processes.
problem Estimating point process models with sensitive personal data raises privacy concerns.
method Proposes differential privacy for event stream data and two optimization algorithms.
result Efficiently estimates Hawkes process models with privacy and utility guarantees.
A large amount of observational data has been accumulated in various fields in recent times, and there is a growing need to estimate the generating processes of these data. A linear non-Gaussian acyclic model (LiNGAM) based on the non-Gaussianity of external influences has been proposed to estimate the data-generating …
A new algorithm splits Gaussian processes for efficient streaming data.
problem Poor scaling of Gaussian processes in streaming data.
method Sequential partitioning of input space and localized Gaussian process fitting.
result The algorithm achieves linear memory complexity and superior time and space complexity.
The paper shows how to efficiently generate large Gaussian process samples with reliability guarantees.
problem Generating large-scale Gaussian process samples efficiently and with reliability.
method Demonstrates scaling data generation to large \(n\) while providing high probability guarantees.
result Efficiently generates large Gaussian process samples with reliability guarantees.
Method identifies regions of maximum dissimilarity in stochastic processes.
problem Comparing local characteristics of two random processes to find periods of maximum dissimilarity.
method Bayesian inference with integrated nested Laplace approximation for stochastic processes.
result Identifies regions of maximum dissimilarity with a certain volume.
A new method for generating synthetic data using posterior distribution learning accelerates inference.
problem Generating high-quality synthetic data requires many discretization steps, which is computationally expensive.
method Learning the posterior distribution of clean data samples given noisy versions, using a scoring rule instead of regression loss.
result Consistently outperforms standard diffusion models at few discretization steps.
Unified framework for deriving generalization bounds in supervised learning.
problem Generalization error bounds in supervised learning.
method Data Processing Inequality PAC-Bayesian framework.
result Unified bounds on binary Kullback-Leibler generalization gap for various divergences.
The aim of process discovery, originating from the area of process mining, is to discover a process model based on business process execution data. A majority of process discovery techniques relies on an event log as an input. An event log is a static source of historical data capturing the execution of a business proc…
We propose a simple method that combines neural networks and Gaussian processes. The proposed method can estimate the uncertainty of outputs and flexibly adjust target functions where training data exist, which are advantages of Gaussian processes. The proposed method can also achieve high generalization performance fo…
Introduces a new stationary GE-process for gold price analysis.
problem Analyzing gold price data with a flexible stationary process.
method Developed a new stationary GE-process with three parameters. Analyzed synthetic and real gold price data.
result Maximum likelihood estimators can be obtained for the unknown parameters.
Social goods, such as healthcare, smart city, and information networks, often produce ordered event data in continuous time. The generative processes of these event data can be very complex, requiring flexible models to capture their dynamics. Temporal point processes offer an elegant framework for modeling event data …
Often in machine learning, data are collected as a combination of multiple conditions, e.g., the voice recordings of multiple persons, each labeled with an ID. How could we build a model that captures the latent information related to these conditions and generalize to a new one with few data? We present a new model ca…
Paper improves neural ODEs for forecasting non-Markovian processes.
problem Forecasting irregularly observed time series with incomplete data.
method Path-dependent Neural Jump ODEs with signature transform.
result Path-dependent NJ-ODE outperforms original framework in non-Markovian data.
Image post-processing is used in clinical-grade ultrasound scanners to improve image quality (e.g., reduce speckle noise and enhance contrast). These post-processing techniques vary across manufacturers and are generally kept proprietary, which presents a challenge for researchers looking to match current clinical-grad…
Simulators enable learning with generalization guarantees in computationally bounded worlds.
problem Generalization in learning theory
method Simulatable processes
result Recovery of PAC learning guarantees with VC dimension
Algorithm ensures demographic parity in regression without sensitive attribute data.
problem Performing regression with demographic parity constraints.
method Post-processing algorithm using accurate estimates and sensitive attribute predictor.
result Generates predictions meeting demographic parity constraint.
Neural Jump ODEs model Itô processes without adversarial training.
problem Generating samples from Itô processes with irregular data.
method Neural Jump ODEs framework for drift and diffusion approximation.
result NJODEs can recover true parameters of Itô processes in the limit.
This paper introduces a multi-output Gaussian process for censored data.
problem Modeling bias in censored data using correlations between multiple outputs.
method Heteroscedastic multi-output Gaussian process with input-dependent noise and variational inference.
result The model better estimates the true process under complex censoring dynamics.
LEAP identifies latent causal variables from temporal data.
problem Recovering time-delayed latent causal variables from general temporal data.
method Proposes LEAP, a framework that extends VAEs with constraints for temporally causal latent processes.
result Successfully identifies temporally causal latent processes from observed variables under various dependency structures.
Multitask Gaussian process regression reduces data generation costs for molecular property prediction.
problem Data bottleneck in training surrogate models for molecular properties.
method Multitask Gaussian process regression over heterogeneous data sources (CC and DFT).
result Predicts at CC-level accuracy with over an order of magnitude reduction in data generation cost.
AutoAIViz visualizes AutoAI's model generation process, improving user trust.
problem Limited transparency in AutoAI systems leads to lack of user understanding and trust.
method Developed and evaluated AutoAIViz, an experimental system that visualizes AutoAI's model generation process.
result AutoAIViz helps users complete data science tasks and increases their understanding, improving trust in AutoAI systems.
New MGCPP model for order flow in financial markets.
problem Modeling order flow dynamics in financial markets.
method Developed MGCPP, proved LLN and FCLTs, applied to real data.
result Validated MGCPP model with real trading data.
We study critera for a pair ({Xn}, {Yn}) of approximating processes which guarantee closeness of moments by generalizing known results for the special case that Yn=Y for all n and Xn converges to Y in probability. This problem especially arises when working with surrogate models, e.g. …