Bayesian model learns complex multivariate dependencies.
problem Learning dependency structures across multiple dimensions.
method Flexible Gaussian process priors and Dirichlet process for structure learning.
result Efficient variational inference for model parameters.
New method measures model risk in dynamic settings with uncertain state processes.
problem Lack of non-parametric approach for dynamic model risk quantification.
method Generalizes relative-entropic approach to dynamic case under f-divergence. result Unified treatment for worst-case risk and f-divergence budget. New graphical criteria for efficient covariate adjustment in non-parametric causal models.
problem Estimating population average treatment effects in observational studies using non-parametric causal graphical models.
method Developed new graphical criteria to determine efficient covariate adjustment sets for estimating treatment effects in non-parametric causal graphical models.
result Graphical criteria for efficient covariate adjustment can be applied in both linear and non-parametric causal models.
Estimates dependent parameters using Markovian dependence with shrinkage.
problem Estimating dependent parameters from a hidden Markov model.
method Developed a novel non-parametric shrinkage algorithm combining Tweedie-based ideas and efficient state estimation.
result Superior performance compared to non-shrinkage methods in hidden Markov models.
SurvMixClust clusters survival data and predicts individual survival curves.
problem Integrating clustering into survival analysis for precision medicine.
method SurvMixClust learns latent representations for clustering and predicts survival functions using a mixture of non-parametric experts.
result SurvMixClust creates balanced clusters with distinct survival curves, outperforming clustering baselines and competing with non-clustering models in predictive accuracy.
Method learns Markov networks from continuous data without distributional assumptions.
problem Learning Markov network structures for continuous data without distributional assumptions.
method Combines non-parametric mutual information estimator with constraint-based algorithm for learning graph structure.
result Shows superior structure learning accuracy compared to competing methods on synthetic data with non-linear dependencies.
New methods using vine copulas improve accuracy of feature dependence in predictive models.
problem Inaccurate feature dependence assumptions in Shapley values lead to incorrect explanations.
method Proposed two new approaches based on vine copulas to model feature dependence.
result Vine copula approaches give more accurate approximations to true Shapley values.
New method tests independence in time series data.
problem Testing independence between time series data.
method Temporal dependence statistic with block permutation.
result Asymptotically valid and universally consistent test for independence.
A new statistical model uses Orlicz-Sobolev spaces with Gaussian weight.
problem Statistical modeling of infinite-dimensional probability measures.
method Affine statistical bundle on Gaussian Orlicz-Sobolev space.
result Provides tools for solving infinite-dimensional evolution problems.
Study evaluates policies in partially observable environments without full model specification.
problem Evaluating policies in partially observable environments without full model specification.
method Developed non-parametric identification and recursive fitted-Q-evaluation algorithm.
result Established finite-sample error bounds for policy value estimation.
New algorithm finds minimal causal models for complex latent variables.
problem Learning causal structure in presence of latent variables and measurement dependencies.
method Graph theoretic edge clique cover problem, non-parametric algorithm.
result Minimality in minimal causal models implies specific properties.
Datasets with hundreds of variables and many missing values are commonplace. In this setting, it is both statistically and computationally challenging to detect true predictive relationships between variables and also to suppress false positives. This paper proposes an approach that combines probabilistic programming, …
Study compares non-parametric models for predicting medical insurance reimbursement delays.
problem Estimating the time-lapse between medical insurance reimbursement.
method Comparative study of four non-parametric regression models (KNNs, SVMs, Decision Trees, Random Forests) using R-squared metric.
result Each model's performance varies with training data size, feature space, and hyperparameters.
New process capability index for non-normal data.
problem Measuring process capability when data does not follow normal distributions.
method Developed a new multivariate non-parametric PCI using Support Vector Data Description (SVDD).
result Demonstrated improved accuracy in process capability measurement for non-normal data.
Distributed Gradient Descent achieves optimal rates in non-parametric regression with linear speed-up.
problem Optimal statistical rates in decentralized non-parametric regression.
method Distributed Gradient Descent with i.i.d. samples and linear speed-up.
result Achieves optimal statistical rates with linear speed-up in the big data regime.
New measure assesses predictive dependence between continuous variables, capturing non-functional relationships.
problem Quantifying the joint dependence between continuous random variables.
method Introduces a novel, fully non-parametric measure bounded [0,1] that assesses predictive accuracy loss.
result The measure captures a wide range of relationships, including non-functional ones, and is interpretable.
Study non-parametric value function estimation from a single path.
problem Estimating value function from a single trajectory in Markov reward processes.
method Kernel-based multi-step temporal difference (TD) estimates, including K-step look-ahead TD and TD(λ). result Non-asymptotic guarantees for TD estimates, capturing interactions between mixing time and model mis-specification.
New neural network models extreme value distributions with preserved shape constraints.
problem Modeling multivariate extreme value distributions with preserved shape constraints.
method d-max-decreasing neural network architecture for non-parametric calibration and generation of MEVs.
result The proposed architecture approximates the dependence structure of MEVs at parametric rate and preserves essential shape constraints.
This work develops a non-parametric test for relational independence in non-i.i.d. data.
problem Testing independence in relational systems where data samples are not i.i.d.
method Kernel mean embedding for relational variables, consistent non-parametric scalable kernel test.
result Empirically validated effectiveness compared to state-of-the-art tests.
Bayesian method models financial time series with non-stationarity and dependency.
problem Discrimination between non-stationarity and long-range dependency in financial time series.
method Adaptive spectral technique using non-parametric Bayesian inference with Reversible Jump Markov Chain Monte Carlo.
result Bayesian method effectively models both long-range dependency and non-stationarity in financial time series.
The paper develops a new method for estimating non-parametric regression functions with spatio-temporal dependencies.
problem Estimating non-parametric regression functions with spatio-temporal dependencies.
method Locally Adaptive Regression Splines (LARS) with ADMM algorithm.
result The method shows superior performance compared to existing techniques.
KQT-EWMA monitors multivariate data streams online with flexible and practical change detection.
problem Online monitoring of multivariate data streams for detecting changes.
method Combines Kernel-QuantTree histogram and EWMA statistic for non-parametric monitoring.
result Controls Average Run Length (ARL0) while achieving comparable detection delays.
A new growth model for dynamic networks using Markovian latent points.
problem Modeling temporal dynamic networks with latent points and distances.
method Markovian latent space dynamic with Euclidean Sphere sampling and connection probabilities based on geodesic distances.
result Theoretical guarantees for non-parametric estimation of the latitude and envelope functions.
A new MFG framework for evolving clusters from Gaussian mixtures.
problem Evolutionary clustering of time-dependent Gaussian mixtures.
method Control-theoretic framework based on Mean Field Games (MFG) with coupled HJB and Fokker-Planck systems.
result MFG dynamics recover classical EM algorithm trajectories with mass conservation.
Near-optimal tests and confidence sequences for non-parametric data.
problem Flexible statistical inference and decision-making with non-parametric data.
method Classic delayed-start normal-mixture sequential probability ratio tests with asymptotic guarantees.
result Asymptotically optimal type-I error and expected rejection time guarantees.
We show that the jumps correlation matrix of a multivariate Hawkes process is related to the Hawkes kernel matrix through a system of Wiener-Hopf integral equations. A Wiener-Hopf argument allows one to prove that this system (in which the kernel matrix is the unknown) possesses a unique causal solution and consequentl…
Develops a non-parametric Dirichlet process method for probabilistic biclustering.
problem Challenges in finding biclusters with strong co-occurrence in rows and columns.
method Dual Dirichlet process mixture models for row and column clustering, with cluster number determined by data.
result Improves bicluster extraction in text mining and gene expression analysis.
A method detects changes in heterogeneous data streams over graph nodes.
problem Detecting changes in data streams from nodes of a graph.
method Online non-parametric method using likelihood-ratio estimation.
result The method accurately identifies change-points in real-world applications.
A new principle for extrapolating regression outside training data.
problem Regression extrapolation when predictions are outside training data range.
method Data-adaptive marginal transformation and simple relationship assumption.
result Progression method offers guarantees on approximation error beyond training data range.
A new MI estimator reduces complexity to linear time, achieving optimal MSE rates.
problem High computational complexity of MI estimators.
method Ensemble Dependency Graph Estimator (EDGE) combining LSH, dependency graphs, and ensemble bias-reduction.
result EDGE achieves optimal computational complexity O(N) and parametric MSE rate O(1/N). New deep learning model uses self-attention to consider entire dataset for predictions.
problem Traditional deep learning models focus on single input datapoints; this model considers the whole dataset.
method Introduces self-attention mechanism to reason about relationships between datapoints.
result Models solve cross-datapoint lookup and complex reasoning tasks.
The strategy of early stopping is a regularization technique based on choosing a stopping time for an iterative algorithm. Focusing on non-parametric regression in a reproducing kernel Hilbert space, we analyze the early stopping strategy for a form of gradient-descent applied to the least-squares loss function. We pro…
Improves normalizing flows by incorporating data dependencies.
problem Current normalizing flow learning assumes independent data, leading to errors.
method Proposes a likelihood objective with dependencies and efficient learning algorithm.
result Improves density estimation and data generation on real-world data.
Most data is multi-dimensional. Discovering whether any subset of dimensions, or subspaces, of such data is significantly correlated is a core task in data mining. To do so, we require a measure that quantifies how correlated a subspace is. For practical use, such a measure should be universal in the sense that it capt…
Neural networks estimate SDEs with jump noise using a Tamed-Milstein scheme.
problem Estimating drift and diffusion functions in SDEs with jump noise.
method Tamed-Milstein scheme with neural networks as non-parametric approximators.
result Flexible estimation of complex nonlinear dynamics in systems with state-dependent noise.
Adapts to high dimensions for estimating conditional moments.
problem Estimation and inference in high-dimensional settings with unknown intrinsic dimension.
method Sub-sampled k-NN Z-estimator, adaptive data-driven sub-sampling. result Estimation error of n−1/(d+2) and asymptotic normality with n1/(d+2) rate. TCMI assesses mutual dependence of continuous variables without parametric assumptions.
problem Estimating mutual information from continuous distributions.
method TCMI extends mutual information to continuous variables using cumulative distributions.
result TCMI facilitates feature selection and ranking of variable sets.
We introduce a copula mixture model to perform dependency-seeking clustering when co-occurring samples from different data sources are available. The model takes advantage of the great flexibility offered by the copulas framework to extend mixtures of Canonical Correlation Analysis to multivariate data with arbitrary c…
New multivariate dependency measure using Gaussian kernel and copula.
problem Measuring dependency between multivariate distributions.
method Gaussian kernel distance to uniform copula, normalization, nonparametric estimate.
result Proposed measure satisfies desirable properties and is compared with existing measures.
A new test for conditional independence adapts to nonlinear dependencies efficiently.
problem Testing conditional independence in nonlinear and high-dimensional data.
method Nearest-neighbor estimator of conditional mutual information combined with local permutation scheme.
result The test reliably simulates null distribution and is better calibrated for non-smooth densities.
Efficient bandit exploration for various distributions without distribution-specific tuning.
problem Optimizing exploration in multi-armed bandit models for different distributions.
method Sub-sampling Duelling Algorithms (SDA) with Random Block sampling for efficient exploration.
result Achieves asymptotically optimal regret for Bernoulli, Gaussian, and Poisson distributions.
DAIF learns fusion structure from data to improve multimodal supervised learning.
problem Tackles the challenge of determining optimal fusion granularity across heterogeneous data sources.
method Combines random matrix theory and non-parametric dependence measures to learn fusion structure directly from data.
result DAIF outperforms state-of-the-art techniques in predicting T-cell differentiation marker protein expression and patient survival.
Estimation of response functions is an important task in dynamic medical imaging. This task arises for example in dynamic renal scintigraphy, where impulse response or retention functions are estimated, or in functional magnetic resonance imaging where hemodynamic response functions are required. These functions can no…
New defense method for non-parametric classifiers robust against adversarial attacks.
problem Lack of robustness in non-parametric classifiers against adversarial attacks.
method Adversarial pruning method to preprocess datasets and a novel attack.
result Adversarial pruning provides a robust defense for non-parametric classifiers.
KL-UCB-switch optimizes bandit strategies for both distribution-dependent and distribution-free performance.
problem Optimizing regret bounds for stochastic bandits.
method Combining MOSS and KL-UCB strategies.
result Achieves both optimal distribution-dependent and distribution-free regret bounds.
A clustering method for multivariate populations with similar dependence structures.
problem Grouping populations with similar dependence structures.
method Orthogonal projection coefficients of density copulas estimated from populations.
result Clusters of populations with similar dependence structures.
Proposes method for eliciting non-parametric joint priors using normalizing flows.
problem Learning complex non-parametric joint priors for model parameters.
method Expert elicitation combined with normalizing flows for generative modeling.
result Framework supports elicitation of both parametric and non-parametric priors.
In data science, it is often required to estimate dependencies between different data sources. These dependencies are typically calculated using Pearson's correlation, distance correlation, and/or mutual information. However, none of these measures satisfy all the Granger's axioms for an "ideal measure". One such ideal…