Spatial and time-dependent data is of interest in many applications. This task is difficult due to its complex spatial dependency, long-range temporal dependency, data non-stationarity, and data heterogeneity. To address these challenges, we propose Forecaster, a graph Transformer architecture. Specifically, we start b…
Deep neural nets can estimate regression with dependent data without the curse of dimensionality.
problem Regression with dependent data and structural assumptions on the regression function.
method Deep recurrent neural network estimate under suitable structural assumptions.
result Deep neural nets can circumvent the curse of dimensionality for regression with dependent data.
Estimating the dependency of variables is a fundamental task in data analysis. Identifying the relevant attributes in databases leads to better data understanding and also improves the performance of learning algorithms, both in terms of runtime and quality. In data streams, dependency monitoring provides key insights …
Survey on learning with graph-dependent data, deriving new generalization bounds.
problem Traditional i.i.d. data assumption fails in many real-life applications.
method Collect and analyze graph-dependent concentration bounds, derive generalization bounds.
result New generalization bounds for graph-dependent data.
Improves normalizing flows by incorporating data dependencies.
problem Current normalizing flow learning assumes independent data, leading to errors.
method Proposes a likelihood objective with dependencies and efficient learning algorithm.
result Improves density estimation and data generation on real-world data.
Proposes a new dependency function for measuring non-linear relationships.
problem Need for a general-purpose measure of dependency between random variables.
method Revision of ideal properties and proposal of a new dependency function.
result Proposes a new dependency function that meets all desired properties.
Estimates binary labels from dependent data using Markov Random Fields.
problem Statistical estimation from dependent data across spatial, temporal, and social domains.
method Modeling dependencies as Markov Random Fields and providing efficient estimation algorithms.
result Statistically efficient estimation rates for Ising models from a single sample.
This research evaluates measures of dependence for financial time-series data.
problem Accurately preparing time series data and selecting an appropriate measure of dependence is challenging.
method Review and establishment of a comprehensive analysis framework for shaping time-series data and evaluating measures of dependence.
result A method, framework, and example for selecting and evaluating a suitable measure of dependence are presented.
Diffusion Transformer captures spatial-temporal dependencies in sequential data.
problem Capturing rich spatial and temporal dependencies in sequential data.
method Established theoretical guarantees for diffusion transformers learning Gaussian process data.
result Spatial-temporal dependencies are captured within attention layers of diffusion transformers.
In data science, it is often required to estimate dependencies between different data sources. These dependencies are typically calculated using Pearson's correlation, distance correlation, and/or mutual information. However, none of these measures satisfy all the Granger's axioms for an "ideal measure". One such ideal…
A new measure of dependence for various data types.
problem Measuring dependence in multivariate, functional, and structured data.
method Combines local normalization with RKHS flexibility.
result Validates the measure's properties and competitive performance.
We present a study of generalization for data-dependent hypothesis sets. We give a general learning guarantee for data-dependent hypothesis sets based on a notion of transductive Rademacher complexity. Our main result is a generalization bound for data-dependent hypothesis sets expressed in terms of a notion of hypothe…
Paper extends learning theory to dependent data with uniform risk bounds.
problem Learning with dependent data sequences.
method Derives uniform risk bounds for dependent data using VC-dimension and Rademacher complexity.
result Standard classification risk bounds hold for dependent data, same as for independent data.
The Probably Approximately Correct (PAC) Bayes framework (McAllester, 1999) can incorporate knowledge about the learning algorithm and (data) distribution through the use of distribution-dependent priors, yielding tighter generalization bounds on data-dependent posteriors. Using this flexibility, however, is difficult,…
We reformulate data-dependent constraints to ensure they are always met with high probability.
problem Ensuring fairness and stability in machine learning models with data-dependent constraints.
method Calibrated reformulation of constraints to guarantee satisfaction with a specified probability.
result Our method guarantees that fairness constraints are met at test time with high probability.
A crucial assumption in most statistical learning theory is that samples are independently and identically distributed (i.i.d.). However, for many real applications, the i.i.d. assumption does not hold. We consider learning problems in which examples are dependent and their dependency relation is characterized by a gra…
The paper shows robustness and generalization are closely connected via data-dependent bounds.
problem Connecting robustness and generalization in machine learning.
method Data-dependent generalization bounds that reduce dependence on covering number and hypothesis space.
result Proves robustness implies generalization, with near-exponential improvements in various situations.
New method uses Winsorized mean estimators for privacy-preserving statistics on dependent data.
problem Privacy-preserving statistics on dependent data with sensitive information.
method Adapting noisy Winsorized mean estimators to handle dependence via log-Sobolev inequalities.
result Asymptotic and finite sample guarantees for item-level and user-level mean estimation similar to \iid{} settings.
The book chapter discusses tail risk analysis for financial data using extreme value statistics.
problem Serial dependence in financial time series complicates tail risk assessment.
method The approach involves unconditional and conditional quantile forecasting.
result Serial dependence impacts multivariate tail dependence.
Bayesian test assesses dependence between mixed data types.
problem Assessing dependence between text, image, and sound data.
method Bayesian kernelised correlation test using Dirichlet process model.
result Demonstrated effectiveness compared to other methods.
Improved convergence for nonconvex optimization with dependent data.
problem Constrained smooth nonconvex optimization with dependent data.
method Stochastic projected gradient methods under a general dependent data sampling scheme.
result Achieved worst-case rate of convergence ildeO(t−1/4) and complexity ildeO(ε−4). Paper studies CLT rates for dependent data in Wasserstein-p distance.
problem CLT rates for multivariate dependent data in Wasserstein-p distance.
method Analyzes locally dependent sequences and geometrically ergodic Markov chains.
result Establishes optimal W1 CLT rates and Wp (p≥2) rates for dependent data. The paper analyzes time-dependent streaming data with biased gradient estimates and proposes improved stochastic optimization methods.
problem Stochastic optimization in a streaming setting with time-dependent and biased gradient estimates.
method Analysis of several first-order methods including SGD, mini-batch SGD, and time-varying mini-batch SGD, along with their Polyak-Ruppert averages.
result Time-varying mini-batch SGD methods can break long- and short-range dependence structures, and biased SGD methods can achieve comparable performance to their unbiased counterparts.
PAC-Bayesian theory applied to data-dependent hypothesis sets yields uniform generalization bounds.
problem Proving uniform generalization bounds for data-dependent hypothesis sets.
method Applying PAC-Bayesian framework on 'random sets' and considering data-dependent hypothesis sets.
result Data-dependent uniform generalization bounds are proven, providing tighter and unified results.
New measures capture tail dependence and non-exchangeability in financial data.
problem Underestimation of tail dependence and inability to capture non-exchangeable tail dependence.
method Tail copulas and novel tail dependence measures (MTCM, ATCM) are proposed.
result Captures non-exchangeable tail dependence and provides analytical forms for various copulas.
Paper introduces data-dependent SSP for private linear and logistic regression.
problem Private linear and logistic regression with better performance.
method Data-dependent sufficient statistic perturbation (SSP) for linear and logistic regression.
result Data-dependent SSP outperforms state-of-the-art methods for linear and logistic regression.
CDSSL improves representation quality by integrating linear and nonlinear dependencies.
problem Scarcity of labeled data and neglect of nonlinear dependencies in SSL.
method CDSSL combines linear correlations and nonlinear dependencies using HSIC in RKHS.
result CDSSL enhances representation quality on diverse benchmarks.
New algorithm achieves data-dependent regret bounds in MDPs with unknown transitions.
problem Achieving best-of-both-worlds guarantees with data-dependent regret bounds in MDPs with unknown transitions.
method Optimistic follow-the-regularized-leader algorithm with new optimistic Q-function estimators and transition bonus.
result First-order, second-order, and path-length bounds with polylog(T) regret in the stochastic regime.
Optimal kernel in KR can be data-dependent, improving model performance.
problem Fixed kernel in KR limits model performance.
method Considered data-dependent kernels for KR, using posterior covariance.
result Data-dependent kernel choice leads to optimal performance.
Conventional sequential learning methods such as Recurrent Neural Networks (RNNs) focus on interactions between consecutive inputs, i.e. first-order Markovian dependency. However, most of sequential data, as seen with videos, have complex temporal dependencies that imply variable-length semantic flows and their composi…
Discovering temporal lagged and inter-dependencies in multivariate time series data is an important task. However, in many real-world applications, such as commercial cloud management, manufacturing predictive maintenance, and portfolios performance analysis, such dependencies can be non-linear and time-variant, which …
Paper generalizes Gaussian universality and CGMT to dependent data, impacting data augmentation in high-dimensional logistic regression.
problem Limitation of Gaussian universality and CGMT in handling dependent data.
method Generalizes Gaussian universality and CGMT to dependent data (block dependence, m-dependence, mixing). Establishes a novel CGMT framework.
result Gaussian universality holds for high-dimensional logistic regression under various types of dependence.
Study improves ERM for heavy-tailed data with dependent inputs.
problem Empirical Risk Minimization with dependent and heavy-tailed data.
method Extending risk bounds for ERM with heavy-tailed, dependent data.
result Established risk bounds for ERM with dependent and heavy-tailed data.
BAM model learns graph structure from data with robustness across linear and non-linear dependencies.
problem Detecting dependencies in datasets for graph structure learning.
method Proposes BAM, a neural network model using structural equation models and Chebyshev polynomials for training, with bilinear attention mechanism.
result Demonstrates robust generalizability and superior performance in graph estimation.
Random Forests adapted for dependent data using GLS.
problem Random Forests' local nature fails for dependent data.
method RF-GLS: GLS loss for correlated data.
result Consistency of RF-GLS under beta-mixing processes.
Temporal data are increasingly prevalent in modern data science. A fundamental question is whether two time series are related or not. Existing approaches often have limitations, such as relying on parametric assumptions, detecting only linear associations, and requiring multiple tests and corrections. While many non-p…
The paper improves PAC-Bayes bounds for data-dependent predictors.
problem Guaranteeing the quality of predictions on unseen examples.
method Basic PAC-Bayes inequality for stochastic kernels, leading to various bounds.
result Validates PAC-Bayes bounds without fixed 'data-free' priors and bounded losses.
Classical causal and statistical inference methods typically assume the observed data consists of independent realizations. However, in many applications this assumption is inappropriate due to a network of dependences between units in the data. Methods for estimating causal effects have been developed in the setting w…
Improves causal graph learning on dependent binary data.
problem Challenges in learning causal graphical models from dependent binary data.
method Decorrelation-based approach using latent utility model and EM-like algorithm.
result Significant improvement in accuracy of causal graph learning.
A new boosting model handles dependent censoring in time-to-event data.
problem Independent censoring assumption leads to biased predictions in time-to-event analysis.
method Clayton-boost, a boosting approach using Clayton copula.
result Clayton-boost outperforms other methods in handling dependent censoring.
New method tests Granger non-causality in panel data with cross-sectional dependencies.
problem Testing Granger non-causality in panel data with cross-sectional dependencies.
method Proposes a new approach to aggregate p-values from panel members to test Granger non-causality, showing lower FDR.
result Our approach discovers true causal relations in panel data, unlike state-of-the-art methods.
Finding relationships between multiple views of data is essential both for exploratory analysis and as pre-processing for predictive tasks. A prominent approach is to apply variants of Canonical Correlation Analysis (CCA), a classical method seeking correlated components between views. The basic CCA is restricted to ma…
Study shows DQN's performance degrades with temporal dependence in data.
problem Temporal dependence in replayed data affects DQN's performance.
method Modelled τ-mixing data, derived risk bounds, and empirical validation. result Temporal dependence leads to a degradation in DQN's performance rate.
Generative model captures complex dependence in financial data.
problem Complex dependence structure in business and financial data.
method Multivariate generative model with heterogeneous and asymmetric tail dependence.
result Novel moment learning algorithm for scalable parameter estimation.
Neural network method estimates covariate-dependent graphical models with statistical guarantees.
problem Estimating graph structure from covariate-dependent data.
method Neural network approach that allows flexible functional dependency on covariates.
result Theoretical PAC guarantees for the method's performance.
Auto-regressive diffusion models improve capturing conditional dependence in data.
problem Vanilla diffusion models struggle to capture important, high-level relationships in real-world data.
method Developed auto-regressive diffusion models to better capture conditional dependence structures.
result AR diffusion models produce samples with a reduced gap in approximating the data conditional distribution.
A new vine copula mixture model improves clustering accuracy for non-Gaussian data.
problem Finite mixture models struggle with asymmetric tail dependencies and non-elliptical clusters.
method Proposes a vine copula mixture model for clustering non-Gaussian data, addressing model selection and parameter estimation.
result Significant improvement in clustering accuracy for data with asymmetric tail dependencies or non-Gaussian margins.
CDPs visualize causal dependencies in AI models.
problem Understanding how AI models depend on data inputs causally.
method Developed Causal Dependence Plots (CDPs) to visualize causal dependencies.
result CDPs show causal changes in predictors and outcomes.