New framework resolves central limit behavior in differential privacy.
problem Choosing appropriate privacy metrics in hypothesis testing.
method Infinitely divisible limit experiments and Le Cam's theory.
result Characterizes all limiting baseline trade-off functions in differential privacy.
We reduce variance in monetization metrics for ranking experiments.
problem Heavy-tailed monetization metrics lead to unreliable conclusions in A/B experiments.
method Post-stratification combined with CUPED.
result Significant reduction in variance and improved decision stability.
Efficient inference method for adaptive experiments with tighter confidence sequences.
problem Efficient inference of Average Treatment Effect in a changing policy sequential experiment.
method Semiparametric efficient inference using Adaptive Augmented Inverse-Probability Weighted estimator and asymptotic confidence sequences.
result Derives tighter confidence sequences for adaptive experiments under data-dependent stopping times.
Paper proposes IRL methods for limited interaction scenarios.
problem Learning with limited teacher interaction.
method Curriculum Inverse Reinforcement Learning (CIRL) and Self-Paced Inverse Reinforcement Learning (SPIRL).
result Training strategies can accelerate learning compared to random or batch methods.
SynthER uses generative models to augment limited RL experience.
problem Limited data for reinforcement learning agents.
method SynthER leverages diffusion models to generate synthetic experience data.
result SynthER significantly improves sample efficiency and training of RL agents.
New study shows limits to classifying brain activity from randomized EEG trials.
problem Classifying human brain activity from image stimuli using EEG is challenging.
method Used randomized trials on a larger dataset (20x) to avoid stimulus-time confound.
result Classification accuracy is marginally above chance and statistically significant.
A budgeted experiment design method for causal structure learning with improved efficiency.
problem Learning causal structure with limited experiments.
method Formulated as an optimization problem, solved using a greedy algorithm with submodularity and accelerated variants.
result Achieves $(1-rac{1}{e})$-approximation of optimal value and significantly reduces the number of interventions.
Develops EBOD model with price limits to explain stock price dynamics.
problem Explains stock price dynamics using empirical order flows.
method Empirical order placement and cancellation processes with price limits.
result Asymmetric price limits cause diverging or vanishing stock prices.
This work characterizes the fundamental limit of network pruning using statistical dimension and convex geometry.
problem The fundamental limit of network pruning is still lacking, especially for deep neural networks.
method Directly imposing sparsity constraint on the loss function and using statistical dimension in convex geometry.
result Characterizes the sharp phase transition point as the fundamental limit of pruning ratio.
New BO methods exploit parallel experiments, reducing search time and improving solution quality.
problem Limitation of Bayesian optimization in exploiting parallel experiments.
method Propose new parallel BO paradigms that exploit the structure of the system to partition the design space.
result Significantly reduce search time and increase probability of finding global solutions.
Study on limits of LLM-based multi-agent planning reliability.
problem Reliability limits of LLM-based multi-agent planning.
method Modeling LLM-based multi-agent architecture as a decision network, showing dominance by centralized Bayes decision maker.
result Optimizing multi-agent directed acyclic graphs under communication budget is equivalent to choosing a constrained experiment.
Transformer RL optimizes A/B testing for time series experiments.
problem Challenges in applying A/B testing to time series experiments, especially with limited history and strong assumptions.
method Transformer reinforcement learning approach that conditions allocation on full history and optimizes MSE without restrictive assumptions.
result Consistently outperforms existing designs in synthetic, simulator, and real-world data.
Optimizes trade execution with reinforcement learning for limit orders.
problem Maximizing revenue in a limit order book with market and limit orders.
method Formulated as a dynamic allocation task, uses multivariate logistic-normal distributions for efficient training.
result Outperforms traditional strategies in simulated environments.
Study efficient resource allocation for detecting extreme values.
problem Efficiently allocate limited resources to detect extreme values in various fields.
method Proposes ExtremeHunter algorithm for sequential resource allocation under limited feedback.
result Demonstrates ExtremeHunter outperforms oracle policy in detecting extreme values.
New estimator reduces bias-variance tradeoff in Markovian interference experiments.
problem Estimating impact of interventions in systems with limited resources.
method Differences-In-Q (DQ) estimator for on-policy policy evaluation.
result DQ estimator has exponentially smaller variance than off-policy methods.
Paper presents a collaborative learning model to improve QoE models without sharing sensitive data.
problem Limited data volume and participant profiles lead to over-fitting and poor generalization of QoE models.
method Round-Robin based Collaborative Machine Learning training without sharing datasets.
result The proposed model outperforms conventional centralized and isolated learning methods.
A new network embedding method using diffusion to overcome limitations of random walks.
problem Limitations of random walk based network embedding methods in fragile sampling and disequilibrium networks.
method Proposes a network diffusion based embedding method that captures both depth and breadth information and uses cascades for global network information.
result The diffusion based models are more robust in fragile sampling and highly imbalanced networks.
Active learning improves SR by proposing experiments in data-limited settings.
problem Efficiently gathering data for symbolic regression with physical constraints.
method Query by committee using the Pareto frontier of equations, with physical constraints.
result Reduces data required for SR and achieves state-of-the-art results.
Meta-RL algorithm improves sample efficiency and adaptability.
problem Challenges in meta-reinforcement learning, especially on-policy experience and task uncertainty.
method Develops an off-policy meta-RL algorithm that probabilistically infers task variables.
result Significantly outperforms prior algorithms in sample efficiency and asymptotic performance.
Improves training GANs by escaping limit cycles.
problem Limit cycling behavior in training GANs.
method Predictive Centripetal Acceleration Algorithm (PCAA) combined with Adam.
result PCAA improves convergence rates and effectively trains GANs.
Optimal experiment design reduces unknown structure in fixed experiments.
problem Learning causal structure from limited experiments.
method Characterizes optimal learning strategy, designs experiments efficiently.
result Proposed algorithm is ρ-approximation for bounded degree graphs.
Study shows algorithms benefit from limited target data with many source domains.
problem Adapting to new domains with scarce labeled target data.
method New family of model selection algorithms.
result Beneficial guarantees in scenarios with limited target data.
Graph-based weather prediction adapted for local models.
problem Applying neural weather prediction to limited area modeling.
method Adapting graph-based Neural Weather Prediction approach to local models.
result Validation of multi-scale hierarchical model extension for Nordic region.
The paper develops a method for self-normalized inference in adaptive experiments.
problem Adaptive experiments require a fixed horizon for ATE estimation, but propensities can change.
method The method uses self-normalized martingale limit theory to estimate ATE.
result The Studentized statistic is asymptotically N(0,1) at the prespecified horizon.
Cohomology fractals illustrate complex 3-manifold properties.
problem Visualizing complex cohomology classes in hyperbolic 3-manifolds.
method Ray-tracing cohomology fractals and proving their distribution.
result Cohomology fractals converge to a distribution on the sphere at infinity.
Unified method for multi-defect microscopy image restoration with limited training data.
problem Challenges in applying deep learning methods due to limited training data for multi-defect microscopy images.
method Two-stage approach: data augmentation with GAN and conditional GAN training.
result Proposed method gives comparable or superior results to existing methods in image quality restoration.
Maximizing mutual information selects simple models from limited data.
problem Selecting simple models from finite and potentially noisy data.
method Prior choice that maximizes mutual information between parameters and predictions.
result The method selects a lower-dimensional effective theory by ignoring poorly constrained parameters.
New insights into experience replay in RL algorithms.
problem Understanding the impact of replay capacity and replay ratio in Q-learning.
method Systematic and extensive analysis of experience replay in Q-learning methods, focusing on replay capacity and replay ratio.
result Greater replay capacity significantly improves performance for certain algorithms, while other techniques offer limited benefit.
Generative model improves ECG classification with limited data.
problem Poor performance of RNNs with limited channel ECG data.
method Generative Seq2Seq model to fill missing data, followed by discriminative learning.
result Generative approach outperforms standard RNNs in disease prediction.
Sparse matrix decomposition identifies key design variables for ICF experiments.
problem Improving predictive capability of ICF simulation codes through better understanding of design inputs and outcomes.
method Sparse Principal Component Analysis (SPCA) and Random Forest (RF) surrogate model.
result Identified clusters of design variables related to physical processes, revealing important variables not previously considered.
Bayesian optimization improved for noisy experiments with constraints.
problem Efficiently optimizing parameters with high noise and constraints.
method Derived expected improvement for noisy batch optimization, developed quasi-Monte Carlo approximation.
result Optimization performance on noisy, constrained problems outperforms existing methods.
New confidence intervals improve treatment effect estimation in randomized experiments.
problem Improving confidence intervals for treatment effects in randomized experiments.
method Systematic exploitation of negative dependence or variance adaptivity.
result Achieved nonasymptotic confidence intervals with the same effective sample size as asymptotic ones.
Translating potential disease biomarkers between multi-species 'omics' experiments is a new direction in biomedical research. The existing methods are limited to simple experimental setups such as basic healthy-diseased comparisons. Most of these methods also require an a priori matching of the variables (e.g., genes o…
A new method combines simple binary classifiers to build complex multiclass classifiers, achieving performance limits in a Gaussian setting.
problem Building a sophisticated multiclass classifier from simple binary decisions.
method Combining O(logK) simple binary classifiers to form a K-class classifier. result Explicit performance bounds across various decoding and dimensional regimes for a stylized Gaussian setting.
The study quantifies and compares aleatoric and epistemic discrimination in ML models.
problem Sources of discrimination in ML models and their impact on performance.
method Quantifying aleatoric and epistemic discrimination using statistical experiments and model accuracy.
result State-of-the-art fairness interventions are effective at removing epistemic discrimination but not aleatoric discrimination in datasets with missing values.
Equal experience fairness in recommender systems reduces bias.
problem Bias in data leads to unfair item recommendations.
method Introduces equal experience fairness notion and optimization framework.
result Mitigates unfairness in recommendations with minimal accuracy loss.
The paper investigates the effectiveness of reusing experience in Deep Q-Learning for FPS environments.
problem The high number of interactions required for reinforcement learning limits its practicality.
method The authors test the effectiveness of applying learning update steps multiple times per environmental step in the VizDoom environment.
result Updating learning steps less frequently than every 4th environmental step does not improve performance and can degrade performance.
RER improves sample complexity by updating in reverse order.
problem Theoretical analysis limits RER's convergence rate.
method Tighter analysis for larger learning rates and longer sequences.
result RER converges faster with larger learning rates and longer sequences.
Framework for systemic risk modeling using jointly exchangeable arrays.
problem Systemic risk in insurance portfolios with interactions.
method Jointly exchangeable arrays, central limit theorems, simulation-based validation.
result Asymptotic approximations for total portfolio losses in large portfolios over long time horizons.
Graph Neural Network improves volatility forecasting for 500 S&P stocks.
problem Forecasting short-term realized volatility in a multivariate setting.
method Graph Transformer Network for Volatility Forecasting.
result Our model outperforms benchmarks on 500 S&P stocks.
Paper derives CLT for Bayesian neural networks trained with variational inference.
problem Analyzing the fluctuation behavior of Bayesian neural networks trained with different variational inference schemes.
method Rigorous derivation of CLT for three variational inference schemes: idealized, Bayes-by-Backprop, and Minimal VI.
result Minimal VI scheme has larger variances but is more computationally efficient.
The paper creates nonparametric confidence bands for band-limited functions.
problem Estimating confidence bands for band-limited functions with finite samples and unknown noise.
method Uses Paley-Wiener reproducing kernel Hilbert spaces and gradient-perturbation methods.
result Non-asymptotic guarantees for confidence regions without assuming a parametric model.
The paper addresses statistical inference issues in adaptive experiments.
problem Statistical inference problems in adaptive experiments.
method Explains and fixes statistical inference issues in adaptive experiments using various methods.
result Various methods to stabilize inferences and recover asymptotic normality.
Study finds real-world datasets contain natural experiments that can improve model performance.
problem Detecting natural experiments in real-world datasets for causal inference.
method Synthetic graph simulation and feature selection based on causal links.
result Real-world datasets contain natural experiments that can be exploited for improved model performance.
Mean field Gaussian inference limits mutual information to regularize neural networks.
problem Understanding and quantifying the regularization effect of mean field Gaussian inference.
method Empirically observed and theoretically quantified mutual information limitation through noise.
result Bounding mutual information between parameters and data effectively regularizes neural networks.
Bayesian optimization improves machine learning system tuning with online and offline experiments.
problem Limited simultaneous experiments in complex policy spaces.
method Augment online field experiments with an offline simulator and apply multi-task Bayesian optimization.
result Substantial gains from including biased offline data in live machine learning systems.
Gaussian Processes improve data interpolation from diverse experiments.
problem Interpolation of sparse and inconsistent datasets from various experiments.
method Used Gaussian Processes (GP) for data interpolation, including uncertainty quantification.
result GPs successfully interpolate data and quantify uncertainties, demonstrating consistency across different sources.
The paper improves confidence regions for band-limited functions using tighter norm bounds and majority voting.
problem Constructing reliable confidence regions for band-limited functions from noisy data.
method Improved norm bounds using Hoeffding's inequality and empirical Bernstein bound, majority voting to aggregate intervals.
result Confidence intervals retain their simultaneous coverage guarantee even when aggregated from random subsamples.