Study uniform consistency in nonparametric mixture models and mixed regression.
problem Uniform consistency in nonparametric mixture models and mixed regression models.
method Construct uniformly consistent estimators under general conditions, develop novel technical tools.
result Prove uniform consistency results for nonparametric mixtures and mixed regression models.
Paper extends nonparametric regression bounds for dependent β-mixing samples.
problem Analyzing error in nonparametric regression with dependent data.
method Extends uniform deviation inequalities from independent to dependent β-mixing samples. result Derives generalization bounds for nonparametric regression with dependent data.
Gradient Boosted Mixed Models estimate mean and variance components for clustered data.
problem Limited flexibility in linear mixed models for complex settings.
method Gradient Boosting extended to mixed models with likelihood-based gradients and flexible base learners.
result Accurate recovery of variance components and improved predictive accuracy.
New method recovers causal DAGs from general environments without strict assumptions.
problem Recovering causal DAGs from real-world data with varying distributions.
method Formalizes desiderata for causal representation learning in general environments, leveraging sufficient change conditions up to third-order derivatives.
result Fully recovers latent DAG and identifies latent variables up to minor indeterminacies under nonparametric mixing.
A new method combines machine learning with mixed-effects models for better repeated measurement analysis.
problem Inference of linear coefficients in partially linear mixed-effects models with complex interactions and high-dimensional variables.
method Double machine learning approach to estimate nonparametrically nonlinear variables, then use standard linear mixed-effects techniques to estimate the linear coefficient.
result The estimated fixed effects coefficient converges at the parametric rate and is semiparametrically efficient.
Paper proposes deep neural networks for nonparametric regression from dependent data.
problem Nonparametric regression from strongly mixing observations.
method Minimum error entropy principle applied to deep neural networks.
result Deep neural networks achieve minimax optimal convergence rates for Gaussian errors.
Develops a deep learning framework for various data types.
problem Handling nonparametric regression and classification across different data types.
method Introduces a general framework with two estimators: NPDNN and SPDNN, based on data satisfying generalized Bernstein-type inequalities.
result Both NPDNN and SPDNN estimators are minimax optimal in many classical settings.
Motivated by problems in data clustering, we establish general conditions under which families of nonparametric mixture models are identifiable, by introducing a novel framework involving clustering overfitted \emph{parametric} (i.e. misspecified) mixture models. These identifiability conditions generalize existing con…
New method for summarizing Bayesian mixture models using sliced Wasserstein distances.
problem Estimating the mixing measure in nonparametric Bayesian mixture models.
method Decision-theoretic approach using sliced Wasserstein distances for Gaussian mixtures.
result Effective estimation of the mixing measure and mixture density.
Statistical machine learning methods, especially nonparametric Bayesian methods, have become increasingly popular to infer clonal population structure of tumors. Here we describe the treeCRP, an extension of the Chinese restaurant process (CRP), a popular construction used in nonparametric mixture models, to infer the …
The paper tackles deep learning from dependent data, achieving optimal performance.
problem Deep learning from strongly mixing observations, especially with regularization and optimality.
method Sparse-penalized regularization for deep neural networks, oracle inequality for expected excess risk.
result Deep neural network estimator achieves minimax optimal rate for nonparametric autoregression.
The study compares parametric and nonparametric models for estimating mean-variance mixtures and finds that nonparametric models perform better.
problem Estimating the distribution of a normal mean-variance mixture under uncertainty.
method Comparison of six parametric mixing laws with a grid nonparametric maximum likelihood estimator, using a paired block bootstrap for score comparison.
result Nonparametric models outperform parametric models in estimating the distribution of a normal mean-variance mixture.
Estimates nonparametric densities from mixed samples.
problem Unmixing convex combinations of nonparametric densities from observed groups.
method Proposes an estimator using topic modeling and U-statistics.
result Rate-optimal estimator for nonparametric density estimation.
Generalized Precision Matrix for scalable estimation of nonparametric Markov networks.
problem Estimating conditional independence structure in general distributions for all data types.
method Generalized Precision Matrix (GPM) for mixed-type variables, regularized score matching framework for scalability.
result Validated theoretical results and demonstrated scalability in various settings.
We introduce 'mixed LICORS', an algorithm for learning nonlinear, high-dimensional dynamics from spatio-temporal data, suitable for both prediction and simulation. Mixed LICORS extends the recent LICORS algorithm (Goerg and Shalizi, 2012) from hard clustering of predictive distributions to a non-parametric, EM-like sof…
New empirical process bounds reveal trade-off between dependence and complexity in nonparametric learning.
problem Understanding generalization in nonparametric learning with temporal dependencies.
method Developed bounds on expected supremum of empirical processes under β/ρ-mixing assumptions. result Achieved rates similar to i.i.d. setting under long-range dependence with complex function classes.
Bayesian nonparametric (BNP) models provide elegant methods for discovering underlying latent features within a data set, but inference in such models can be slow. We exploit the fact that completely random measures, which commonly used models like the Dirichlet process and the beta-Bernoulli process can be expressed a…
Improved density estimation for mixed discrete-continuous data.
problem Inconsistent density estimation for mixtures of continuous and discrete data.
method Modification of existing nonparametric density estimation methods to handle mixed discrete-continuous data.
result Improved consistency and empirical performance for mixed discrete-continuous data.
Partition Tree estimates conditional densities for mixed continuous and categorical variables.
problem Estimating conditional densities for mixed data types.
method Tree-based framework modeling conditional distributions as piecewise-constant densities on adaptive partitions, minimizing conditional negative log-likelihood.
result Improved probabilistic prediction compared to CART-style trees and state-of-the-art methods.
We introduce the nonparametric metadata dependent relational (NMDR) model, a Bayesian nonparametric stochastic block model for network data. The NMDR allows the entities associated with each node to have mixed membership in an unbounded collection of latent communities. Learned regression models allow these memberships…
Paper addresses covariate shift in deep learning regression models.
problem Covariate shift in dependent data from different distributions.
method Sparse-penalized deep neural network (SPDNN) estimator for nonparametric regression.
result Adaptive convergence rates for quantile and Huber regression.
Stochastic volatility modelling of financial processes has become increasingly popular. The proposed models usually contain a stationary volatility process. We will motivate and review several nonparametric methods for estimation of the density of the volatility process. Both models based on discretely sampled continuo…
GBMixed boosts mixed models for clustered data, estimating mean and variance flexibly.
problem Flexible estimation of mean and variance components in clustered data.
method Gradient Boosting framework for linear mixed models with likelihood-based gradients.
result GBMixed accurately recovers complex nonlinear fixed effects and covariances.
New ICA method for sources with mixed spectra.
problem Inaccurate separation of sources with temporal autocorrelations and mixed spectra.
method Estimates spectral density functions and line spectra using cubic splines and indicator functions, then maximizes the Whittle likelihood function.
result Outperforms existing ICA methods in simulations and EEG data applications.
GP-MRO discovers robust mixed strategies for unknown objectives.
problem Optimizing unknown objectives against worst-case uncertain parameters.
method Sequential learning from noisy point evaluations, combining online learning and Gaussian processes.
result GP-MRO finds robust mixed strategies that significantly improve performance over deterministic strategies.
Study nonparametric estimator for Markov chain transition matrices in offline setting.
problem Estimating transition matrices of finite controlled Markov chains from logged data.
method Developed sample complexity bounds and conditions for minimaxity.
result Achieving certain statistical risk requires balancing mixing properties and sample size.
Increasingly complex datasets pose a number of challenges for Bayesian inference. Conventional posterior sampling based on Markov chain Monte Carlo can be too computationally intensive, is serial in nature and mixes poorly between posterior modes. Further, all models are misspecified, which brings into question the val…
Bayesian approach learns nonparametric mixture components from heterogeneous data.
problem Realistic modeling of heterogeneous data populations with nonparametric mixture components.
method Bayesian nonparametric modeling using Dirichlet process mixture priors.
result Posterior contraction rates for component densities are nearly polynomial, improving over deconvolution methods.
Study counterfactuals in combinatorial choice using a representative agent model.
problem Analyzing decision-making from aggregated binary polytope data.
method Nonparametric approach based on a representative agent model, solving polynomial and mixed-integer convex programs.
result Developed a method for counterfactual prediction that works even under model misspecification.
Paper tackles robust deep learning from weakly dependent data with unbounded loss and input.
problem Tackles robust deep learning from weakly dependent data with unbounded loss and input.
method Establishes non-asymptotic bounds for expected excess risk under strong mixing and ψ-weak dependence assumptions. result Derives a relationship between bounds and r, and shows convergence rate close to i.i.d. results for r=∞. Sparse-penalized deep neural networks improve performance in weakly dependent processes.
problem Nonparametric regression and classification under weak dependence.
method Sparse-penalized deep neural networks with oracle inequalities and convergence rates established.
result The proposed estimators outperform non-penalized ones in simulations.
Unified Skew-Gaussian process framework for various regression and classification tasks.
problem Handling multiple types of regression and classification problems.
method Generalization of Skew-Gaussian processes to handle various types of data and likelihoods.
result Closed-form posterior distributions for multiple tasks.
Study identifies causal relationships without direct supervision from unknown interventions.
problem Identify causal relationships from unknown interventions without direct supervision.
method General nonparametric setting with multiple datasets from unknown interventions.
result Identify ground truth latents and causal graph up to ambiguities.
Study shows DQN's performance degrades with temporal dependence in data.
problem Temporal dependence in replayed data affects DQN's performance.
method Modelled τ-mixing data, derived risk bounds, and empirical validation. result Temporal dependence leads to a degradation in DQN's performance rate.
We describe a nonparametric topic model for labeled data. The model uses a mixture of random measures (MRM) as a base distribution of the Dirichlet process (DP) of the HDP framework, so we call it the DP-MRM. To model labeled data, we define a DP distributed random measure for each label, and the resulting model genera…
In this paper we present a nonparametric method for extending functional regression methodology to the situation where more than one functional covariate is used to predict a functional response. Borrowing the idea from Kadri et al. (2010a), the method, which support mixed discrete and continuous explanatory variables,…
We present a mixed multinomial logit (MNL) model, which leverages the truncated stick-breaking process representation of the Dirichlet process as a flexible nonparametric mixing distribution. The proposed model is a Dirichlet process mixture model and accommodates discrete representations of heterogeneity, like a laten…
The problem of developing binary classifiers from positive and unlabeled data is often encountered in machine learning. A common requirement in this setting is to approximate posterior probabilities of positive and negative classes for a previously unseen data point. This problem can be decomposed into two steps: (i) t…
The study evaluates the performance of ANNs in financial forecasting.
problem Mixed evidence on the predictive performance of ANNs for financial time-series data.
method Proposes a flexible nonparametric model and compares its performance to other estimators.
result The proposed model shows better performance than basic benchmarks in estimating Value-at-Risk.
While all kinds of mixed data -from personal data, over panel and scientific data, to public and commercial data- are collected and stored, building probabilistic graphical models for these hybrid domains becomes more difficult. Users spend significant amounts of time in identifying the parametric form of the random va…
Perturbation theory improves nonparametric instrumental variable estimation accuracy.
problem Improving nonparametric instrumental variable estimation accuracy in high-dimensional settings.
method Perturbative approach based on physics perturbation theory, extending kernel ridge methods with higher-order corrections.
result First-order perturbative corrections reduce prediction error by up to 99% in high-dimensional ill-defined cases.
Given a heterogeneous time-series sample, the objective is to find points in time (called change points) where the probability distribution generating the data has changed. The data are assumed to have been generated by arbitrary unknown stationary ergodic distributions. No modelling, independence or mixing assumptions…
We describe an adaptation of the simulated annealing algorithm to nonparametric clustering and related probabilistic models. This new algorithm learns nonparametric latent structure over a growing and constantly churning subsample of training data, where the portion of data subsampled can be interpreted as the inverse …
The hierarchical Dirichlet process (HDP) has become an important Bayesian nonparametric model for grouped data, such as document collections. The HDP is used to construct a flexible mixed-membership model where the number of components is determined by the data. As for most Bayesian nonparametric models, exact posterio…
dcFCI discovers causal relationships robustly under latent confounding and mixed data.
problem Causal discovery under latent confounding and unfaithfulness.
method dcFCI integrates a new score to assess PAG compatibility, guided by FCI search.
result Significantly outperforms state-of-the-art methods in small and heterogeneous datasets.
Semiparametric Bayesian networks combine parametric and nonparametric models for flexible data analysis.
problem Combining the advantages of parametric and nonparametric models for flexible data analysis.
method Semiparametric Bayesian networks combining parametric and nonparametric conditional probability distributions. Modifications of two algorithms for structure learning from data.
result Accurately learns the combination of parametric and nonparametric components, comparable to state-of-the-art methods.
The paper improves Fisher-Pitman tests for Poisson mixtures, detecting autism-related genes.
problem Detecting differentially expressed genes between autism and control subjects.
method Nonparametric Poisson mixtures and Fisher-Pitman permutation tests.
result The tests reveal genes missed by common methods, demonstrating rate optimality.
New test detects differences in heterogeneous datasets.
problem Detecting differences between two samples with unknown heterogeneity.
method Developed a nonparametric testing procedure that handles latent heterogeneity through a composite null.
result The test accurately detects differences in the presence of unknown heterogeneity.