Empirical mode modeling improves state-space analysis of noisy data.
problem Analyzing nonlinear systems with noisy data.
method Combining empirical mode decomposition with empirical dynamic modeling.
result Empirical mode modeling enhances state-space representations in noisy data.
Empirical moment matrix reveals properties of point clouds.
problem Uncovering properties of point clouds, especially those with singular support.
method Combining statistics, real algebraic geometry, and approximation theory.
result The empirical moment matrix provides insights into data analysis.
Bayesian Empirical Bayes extends EB to complex structures using probabilistic symmetry.
problem Improving simultaneous inference in complex settings like arrays and graphs.
method Generalized empirical Bayes approach based on probabilistic symmetry.
result BEB outperforms existing methods in denoising arrays and spatial data.
A new DP algorithm for weighted ERM protects sensitive data in predictive models.
problem Protecting sensitive personal information in predictive models trained via ERM.
method Proposes the first differentially private algorithm for weighted ERM with formal privacy guarantees.
result Demonstrates strong DP guarantees while maintaining robust performance in real-world data.
Solves empirical risk minimization for relational data using graph sampling.
problem Empirical risk minimization for relational data.
method Graph sampling theory, stochastic gradient descent, automatic differentiation.
result Automatic unbiased stochastic gradients for relational data.
Paper validates ABM using stylized financial facts.
problem Validate ABM-generated financial data against real-world data.
method Compare ABM results with stylized financial facts.
result Model successfully replicates stylized financial facts.
We introduce a new discriminant analysis method (Empirical Discriminant Analysis or EDA) for binary classification in machine learning. Given a dataset of feature vectors, this method defines an empirical feature map transforming the training and test data into new data with components having Gaussian empirical distrib…
Raising statistical hurdles may not be justified due to data bias.
problem Data bias leads to unobserved results that weaken identification of revised hurdles.
method Theoretical and empirical analysis of statistical hurdles and data bias.
result Statistics targeting only published findings can be strongly identified.
Estimates policy value from off-policy data in contextual bandits.
problem Estimating policy value from limited off-policy data in contextual bandits.
method Empirical likelihood techniques for optimization.
result Improves over previous methods in finite sample regimes.
SBIC learns patterns in imbalanced datasets using empirical similarity and synthetic data.
problem Classification failure in imbalanced datasets.
method SBIC uses an empirical similarity function and absent data to optimize weights and find minority class data points.
result SBIC outperforms other classification techniques for imbalanced datasets.
Combines expert knowledge and data for efficient probability distribution inference.
problem Inferring discrete probability distributions using limited data and expert knowledge.
method A novel estimator that weights expert knowledge and empirical data.
result The proposed estimator is always more efficient than either expert or data alone.
Spectral regularization improves learning over combinatorial spaces with limited data.
problem Learning pseudo-Boolean functions with scarce labeled data.
method Regularizing the spectral representation of learned functions using the L_1 norm.
result Regularization allows for data-frugal learning and achieves statistically optimal generalization performance.
A method for classification using pairwise similarities and unlabeled data.
problem Handling pairwise similarities and unlabeled data for classification.
method Empirical risk minimization approach to create an unbiased risk estimator.
result Derives an unbiased risk estimator for handling both similarities and unlabeled data.
Paper tackles heavy-tailed data without finite variance, proposing robust risk minimization.
problem Empirical risk minimization under heavy-tailed data with finite p-th moment. method Minimizes risk values robustly estimated via Catoni's method, using generalized generic chaining.
result Shows better performance of optimizer based on empirical risks via Catoni-style estimation.
Unified approach to private statistics from empirical to population data.
problem Divided focus on empirical vs population statistics in private statistics.
method Unified methods for both types of statistics.
result Methods for empirical statistics can be applied to population statistics.
ROI-driven data analytics guides investment in empirical data analysis.
problem Determining the optimal depth and breadth of data analytics.
method Conceptual framework validated through empirical studies focusing on dependency extraction in Mozilla Firefox project.
result ROI-driven data analytics helps avoid over-analyzing empirical data.
Empirical study finds variance swap rate is affine in spot variance for S&P500 data.
problem Investigating the relationship between variance swap rate and spot variance.
method Empirical analysis using S&P500 data from 2006-2018, testing different models.
result Affine relationship between variance swap rate and spot variance is supported.
Paper proposes an unbiased classifier from triplet comparison data.
problem Learning a classifier from triplet comparison data.
method Empirical risk minimization framework with an unbiased estimator.
result The proposed method achieves better performance than baseline methods.
Empirical Gaussian Processes learn flexible priors from data.
problem Limited effectiveness of standard Gaussian process kernels.
method Estimate mean and covariance functions empirically from data.
result Empirical GPs converge to closest GP to real data generating process.
New empirical PAC-Bayes bound for Markov chains with finite state space.
problem Lack of empirical bounds for Markov chains with temporal dependence.
method Proved a new PAC-Bayes bound for Markov chains, providing an empirical pseudo-spectral gap.
result First fully empirical PAC-Bayes bound for Markov chains with finite state space.
Empirical median performs well in estimating location with varying scales.
problem Estimating location with varying scales in data.
method Analysis of empirical median as an estimator.
result Matching upper and lower bounds on estimation error.
Data pruning algorithms struggle in high compression regimes, as shown by theoretical and empirical studies.
problem Limitations of score-based data pruning algorithms in high compression regimes.
method Theoretical and empirical analysis of score-based data pruning algorithms.
result Score-based data pruning algorithms fail in high compression regimes due to 'No Free Lunch' theorems.
A heuristic framework tests the multi-manifold hypothesis in empirical data.
problem Overestimation of parameters in global linear models.
method Heuristic multiscale framework using spline-interpolated manifolds.
result Validates the multi-manifold hypothesis in empirical data.
We prove time series data forms a Kolmogorov space with hidden dimensions.
problem Understanding the structure of time series data.
method Defining cyclic coordinates and spinor fields in time series data.
result Time series data has hidden eight dimensions.
This paper revisits the distribution gap between clean and augmented data in deep learning.
problem The distribution gap between clean and augmented data in deep learning models.
method Analytical perspective using empirical risk and generalization error, highlighting data augmentation as regularization.
result Data augmentation significantly reduces generalization error but slightly increases empirical risk.
This paper tightens the law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.
problem Developing nonasymptotic concentration bounds for empirical KL_inf with optimal constants and rates.
method Presenting a tight law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.
result A tight law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.
Solves memorization in diffusion models for manifold data.
problem Memorization effect in diffusion models for manifold data.
method Inertia update at the end of empirical diffusion simulation.
result Approximates true data distribution on a C2 manifold. Mixup improves model accuracy and calibration through data transformation and random perturbation.
problem Improving model accuracy and calibration in machine learning.
method Interprets Mixup as empirical risk minimization with data transformation and random perturbation.
result Mixup induces multiple known regularization schemes that prevent overfitting and overconfident predictions.
New method for efficient inference in large datasets.
problem Statistical inference in massive datasets.
method Combines divide-and-conquer method and empirical likelihood.
result Reduces computation burden and demonstrates effectiveness.
This paper improves random feature sampling using empirical leverage scores.
problem Optimizing the number of features for kernel approximation and supervised learning.
method Uses empirical leverage scores to optimize feature sampling.
result Empirical sampling of random features using leverage scores outperforms vanilla Monte Carlo sampling.
Introduces foundation priors for using model-generated data in empirical research.
problem Using model-generated data as real observations in empirical research.
method Introduces foundation priors as an exponential-tilted, generalized Bayesian update of the user's primitive prior.
result Synthetic data reflects both model patterns and user's priors, enabling principled use in empirical work.
This paper evaluates how different imputation methods affect predictive models.
problem The impact of different imputation methods on predictive models' performance.
method Systematic evaluation of various imputation methods for different data sets and machine learning algorithms.
result Recommendation of a general method for empirical benchmarking of imputation methods.
Study inequality measures in wealth exchange models and compare with empirical data.
problem Analyzing inequality in wealth distribution models.
method Calculated Gini index and k-index, found bounds, and computed exact quantities for specific distributions.
result Found lower and upper bounds for inequality indices and discussed model efficiencies.
Reweighting improves risk bounds in certain data regions.
problem Improving risk bounds in classification and heteroscedastic regression.
method Weighted empirical risk minimization with a data-dependent weight function.
result A weighted ERM estimator can achieve superior performance in specific sub-regions.
Developed accurate empirical potentials for Si:H nanowires using multi-fidelity Gaussian process.
problem Accurate modeling of Si:H nanowires using fast but inaccurate empirical potentials and slow but accurate first-principle calculations.
method Employed multi-fidelity Gaussian process regression to integrate low-fidelity empirical potential data with high-fidelity first-principle calculations.
result Demonstrated the accuracy of developed empirical potentials for Si:H nanowires.
New approach avoids excess empirical risk in domain generalization.
problem Learning models that generalize to unseen distributions from diverse data sets.
method Minimizes penalty under constraint of optimal empirical risk, leveraging rate-distortion theory.
result Significant improvements in domain generalization performance across multiple methods.
New method uses unlabeled data to improve generalization bounds for deep learning.
problem Vacuous guarantees and shrinking holdout sets for overparameterized models.
method Augmenting labeled training set with unlabeled data and training as usual.
result Proves tight upper bounds on true risk for 0-1 empirical risk minimization.
The paper examines the tilted empirical risk's generalization and robustness under negative tilt.
problem The generalization error of machine learning algorithms under negative tilt.
method Uniform and information-theoretic bounds on the tilted generalization error under negative tilt.
result The tilted empirical risk's generalization error has a convergence rate of \(O(n^{-ε/(1+ε)})\).
The MEM method uses data-driven priors for linear inverse problems, proving convergence and estimating differences.
problem Linear inverse problems with approximate priors.
method Maximum Entropy on the Mean (MEM) method with data-driven priors.
result Empirical mean convergence and estimates for prior differences based on epigraphical distance.
ECOD detects outliers without parameters, fast and simple.
problem Detecting outliers in large, high-dimensional datasets efficiently and interpretably.
method ECOD estimates empirical cumulative distribution functions per dimension, then computes tail probabilities and outlier scores.
result ECOD outperforms state-of-the-art methods in accuracy, efficiency, and scalability.
The study evaluates various jump tests for high-frequency financial data.
problem Choosing the most effective jump test for high-frequency financial data.
method An extensive evaluation of multiple alternative tests in various scenarios.
result Guidelines for choosing the most suitable test for different settings.
Improved eigenvalue distribution method for financial data.
problem Noise and complexity in financial markets.
method Matrix H theory, hierarchical structure, informational cascade.
result Captures a larger fraction of data variance in financial markets.
EB-PCA reduces noise in high-dimensional PCA by estimating a joint prior distribution.
problem High-dimensional PCA noise in samples comparable to or larger than data.
method Empirical Bayes PCA using Kiefer-Wolfowitz MLE, random matrix theory, and AMP algorithm.
result EB-PCA achieves Bayes-optimal accuracy in spiked models and significantly improves over PCA in simulations and real data.
The paper proves deep learning can be robust with certain loss functions.
problem The robustness of deep learning models under flawed data.
method Empirical-risk minimization with unbounded, Lipschitz-continuous loss functions.
result These loss functions provide efficient prediction under minimal data assumptions.
A neural network method estimates densities from characteristic functions.
problem Estimating fixed-horizon probability densities from empirical characteristic functions.
method Data-driven Fourier-mixture neural-network method trained in Fourier space.
result Competitive performance and clear gains on heavy-tailed targets.
Develops methods to improve demand counterfactuals from imperfect proxies.
problem Imperfect proxies in demand models lead to biased counterfactuals and invalid inference.
method Practical toolkit for market-level and individual data, requiring minimal computation.
result Improves substitution prediction and counterfactual performance.
Flexible estimator synthesizes noisy experiments and covariates for optimal effect estimation.
problem Simultaneous analysis of many noisy experiments with rich covariate information.
method Plug-in empirical Bayes estimator that synthesizes noisy experimental results and covariates.
result Within a constant factor of minimax for a simple data-generating model, and robust convergence guarantees hold under generality.
Develops EB for implicit likelihoods using simulators.
problem Traditional EB assumes tractable likelihoods, SBEB handles implicit likelihoods.
method Simulation-based empirical Bayes (SBEB) connects nonparametric EB to SBI, iteratively refining EB estimates.
result SBEB improves accuracy over SBI with fixed priors.