SEARNN improves RNN training by incorporating global-local losses.
problem RNNs trained with MLE fail to exploit structured losses and suffer from exposure bias.
method SEARNN introduces global-local losses through test-alike search space exploration.
result SEARNN outperforms MLE on OCR, spelling correction, and machine translation tasks.
Every locally compact local group is locally isomorphic to a topological group.
A new TwinGP framework for efficient large-scale GP modeling.
problem Efficiently modeling large-scale Gaussian processes with computational constraints.
method Combines global and local approximations using a subset-of-data approach.
result TwinGP framework performs on par or better than state-of-the-art methods at a fraction of the computational cost.
Proposes a new method to control FDR using frequentist-assisted horseshoe for high-dimensional testing.
problem Designing tests with frequentist false discovery rate control using horseshoe prior.
method Frequentist-assisted horseshoe procedure for high-dimensional normal means testing.
result Consistently achieves robust finite-sample FDR control in various sparse cases.
Optimizes edge coloring in graph bundling for better edge differentiation.
problem Difficulty in identifying origins and destinations of individual edges in strongly bundled graphs.
method Optimizes edge coloring based on pairwise edge strength and origin-destination dissimilarity, solving a nonlinear optimization problem.
result Peacock bundles enhance graph layout comprehensibility with edge differentiation.
The paper revisits and improves on a Bayesian relevance vector machine method for small sample sizes.
problem Statistical modeling with small sample sizes relative to the number of covariates.
method Introduces a new class of global-local priors and provides theoretical properties.
result Results on posterior consistency and contraction rates are provided.
New framework optimizes complex systems decisions via simulation.
problem Optimizing strategic, tactical, and operational decisions in complex systems.
method Global-local metamodel assisted two-stage optimization via simulation.
result Framework efficiently searches for optimal decisions with unknown objective.
Bayesian method clusters data and selects variables with shrinkage priors.
problem Sparse convex clustering with limited data accuracy issues.
method Bayesian approach using global-local shrinkage priors and Gibbs sampling.
result Improved estimation accuracy in sparse convex clustering.
ETC improves Transformer models for long and structured inputs.
problem Scaling input length and encoding structured inputs in Transformers.
method Introduces global-local attention, relative position encodings, and CPC pre-training.
result Achieves state-of-the-art results on four natural language datasets.
Proposes a global model from local sparsity assumptions in inverse problems.
problem Inability of traditional sparse modeling to model entire global signals optimally.
method Constructs a global model from local sparsity assumptions using constrained underdetermined systems and ADMM optimization.
result Shows unique and stable recovery conditions for global signal representation.
We prove a global local rigidity result for character varieties of 3-manifolds into SL2. Given a 3-manifold with toric boundary M satisfying some technical hypotheses, we prove that all but a finite number of its Dehn fillings Mp/q are globally locally rigid in the following sense: every irreducible repr…
TopoGeoScore selects robust checkpoints using only source-domain representations.
problem Selecting robust checkpoints without target-domain labels or samples.
method Constructs class-conditional mutual k-nearest-neighbour graphs and extracts three interpretable signals.
result Source representations contain measurable global-local-topological evidence of robustness.
Model for dynamic relational data with regime changes.
problem Handling abrupt changes in dynamic relational data.
method Factorized fusion shrinkage model with global-local shrinkage priors.
result Posterior distribution attains minimax optimal rate up to logarithmic factors.
Horseshoe priors improve small area estimation by borrowing strength globally but locally.
problem Improving precision of small area estimators through global-local borrowing of strength.
method Developed a tail-robust horseshoe model for Fay-Herriot small area estimation, using heteroscedastic Tweedie identity and regular variation theory.
result The horseshoe model outperforms structured Gaussian smoothing on strongly spatial data, identifying exceptional areas that smoothing suppresses.
HS-MoE selects sparse experts using adaptive priors and data-adaptive gating.
problem Sparse expert selection in mixture-of-experts architectures.
method Combines horseshoe prior with input-dependent gating for data-adaptive sparsity.
result Data-adaptive sparsity in expert usage.
Sparse Bayesian model improves covariance estimation in high dimensions.
problem Curse of dimensionality in dynamic covariance estimation.
method Latent time-varying stochastic factors with global-local shrinkage prior.
result Precise correlation estimates, strong minimum variance portfolio performance, superior forecasting accuracy.
The study improves equity return forecasts using shrinkage priors and heavy-tailed distributions.
problem Improving equity return forecasting accuracy using Bayesian econometric models.
method Flexible Bayesian state space model with global-local shrinkage priors and heavy-tailed innovations.
result Several variants of the proposed model outperform traditional methods in forecasting accuracy.
New bounds for private learning of high-dimensional Gaussian distributions.
problem Learning high-dimensional Gaussian distributions under differential privacy constraints.
method Analytic tools for constructing global covers from local covers, modified hypothesis selection techniques.
result Near-optimal sample complexity bounds for general Gaussians, conjectured to be near-optimal in the general case.
Bayesian tree ensemble model for estimating treatment effects in high-dimensional survival data.
problem Estimating heterogeneous treatment effects in censored survival data with many covariates.
method Developed a Bayesian tree ensemble model with a horseshoe prior for adaptive shrinkage.
result Accurately estimates treatment effects in high-dimensional covariate spaces and non-linear functions.
A new BO method adapts hyperparameters online and uses a novel kernel for global and local optimization.
problem Expensive black-box optimization problems.
method Online length-scale adaption, mixed-global-local kernel, and adaptive hyperparameters.
result The proposed method outperforms state-of-the-art BO methods on global optimization benchmarks.
TWIST algorithm detects communities in multi-layer networks with tensor decomposition.
problem Community detection in multi-layer networks with multiple node-modality relationships.
method Tensor-based TWIST algorithm for global/local node and layer memberships.
result Accurate community detection with small misclassification error as network size increases.
New PCA method handles multiple datasets and detects sparse patterns robustly.
problem Handling multi-source data with sparse and outlier-robust PCA.
method Developed a regularization problem with a penalty for structured sparsity and outlier resistance.
result The method detects global and local patterns across multiple data sources robustly.
We derive a new radial link for binary classification under shared elliptical distributions.
problem Binary classification under shared-generator elliptical class-conditional distributions.
method We derive the Bayes radial-link family from the within-class radius law and estimate it by a finite fractional-power stochastic-polynomial projection.
result The derived link is asymptotically Bayes-optimal and significantly better than QDA on various benchmarks.
Proposes integrating global and local entropy for more reliable LLMs.
problem Uncertainty in large language models (LLMs) leads to unreliable predictions.
method Measures global uncertainty from hidden-state matrices and local uncertainty from tokens, combining them via a multiplicative gate.
result Global-Local Uncertainty (GLU) outperforms unsupervised baselines across multiple models and benchmarks.
Horseshoe regularization improves machine learning in complex models.
problem Improving machine learning performance in nonlinear and non-Gaussian models.
method Horseshoe regularization for complex models including deep neural networks.
result Demonstrates the versatility of horseshoe regularization in machine learning.
Hybrid model combines deep learning and classical methods for forecasting time series.
problem Forecasting large collections of similar time series is challenging and complex.
method Proposes a hybrid model integrating deep neural networks and classical time series models.
result Demonstrates improved data efficiency, accuracy, and computational complexity.
Meta-GLAR combines global deep representations with local adaptation for improved forecasting accuracy.
problem Joint learning from related time series boosts accuracy but fails for out-of-sample forecasting.
method Meta-GLAR uses a meta-learning approach to adapt RNN representations for each time series.
result Meta-GLAR outperforms state-of-the-art methods in out-of-sample forecasting accuracy.
Bayesian interpretation explains double descent in deep learning models.
problem Understanding the risk function behavior of over-parameterized models.
method Bayesian model selection, Dickey-Savage ratio, ridge regression, global-local shrinkage.
result Double descent phenomenon explained through Bayesian interpretation.
Dual explanation method using convex hulls and example-based vectors.
problem Local and global explanation of complex models.
method Dual representation of instances as convex combinations, generating new dual dataset, training linear surrogate model, computing feature importance.
result Effective example-based and local/global explanation of complex models.
The paper improves Bayesian precision matrix estimation for high-dimensional sparse data.
problem Estimating sparse precision matrices in high-dimensional settings.
method Tempered posterior with fully specified horseshoe prior.
result Concentration results and theoretical oracle inequality for posterior.
Paper addresses offline policy evaluation in RL, achieving near-optimal bounds for various policy classes.
problem Evaluate all policies in a class simultaneously for offline RL.
method Uniform convergence in OPE for various policy classes, achieving optimal episode complexity.
result Achieves optimal episode complexity of O(H^3/d_mε^2) for identifying ε-optimal policies.
This paper deals with chain graphs under the Andersson-Madigan-Perlman (AMP) interpretation. In particular, we present a constraint based algorithm for learning an AMP chain graph a given probability distribution is faithful to. Moreover, we show that the extension of Meek's conjecture to AMP chain graphs does not hold…
Proposes a tail-adaptive shrinkage method for robust sparse estimation.
problem Robust Bayesian methods for high-dimensional regression under diverse sparse regimes.
method Global-local-tail (GLT) Gaussian mixture distribution with tail-adaptive shrinkage.
result GLT posterior contracts at minimax optimal rate for sparse normal mean models.
This paper proposes a new method to improve VI approximations by capturing dependence between blocks using vector copulas.
problem Improving variational inference accuracy for complex models with challenging posteriors.
method Using vector copulas to model dependence between multivariate blocks, with learnable transport maps for flexible marginals.
result The proposed method produces more accurate posterior approximations than existing methods at limited computational cost.
Meta-learning framework improves explainability of GNNs.
problem Improving explainability of graph neural networks.
method Meta-learning framework to steer GNN training towards interpretable minima.
result Models are easier to explain by different algorithms without sacrificing accuracy.
New examples show limits of physical link isotopies.
problem Limits of physical link isotopies in thick and thickly embedded links.
method Construction of specific examples of thick and thickly embedded links.
result Explicit examples of links that cannot be split through thick homotopies.
DeepGLO forecasts high-dimensional time series by combining global and local models.
problem Forecasting high-dimensional time series with global patterns and local calibration.
method Hybrid model combining global matrix factorization and local temporal networks.
result DeepGLO outperforms state-of-the-art approaches by more than 25% in WAPE.
Improved credit scoring model with explainability.
problem Making financial decisions based on loan applications.
method Extreme Gradient Boosting (XGBoost) model with 360-degree explanation framework.
result Model achieves state-of-the-art performance and provides understandable explanations.
BaGGLS models biological interactions using Bayesian shrinkage for interpretability.
problem Interpreting complex interactions in high-dimensional biological data.
method Bayesian group global-local shrinkage prior with variational approximation.
result BaGGLS outperforms other methods in interaction detection and scalability.
A new framework reduces RL training cost by optimizing hyper-parameters.
problem High sampling cost in RL due to complex hyper-parameter tuning.
method Proposes a 'reinforcement on reinforcement' (RoR) architecture to decompose tasks into two layers of RL.
result The proposed framework achieves up to 56% expected sampling cost saving.
Proposes ESCA model to analyze mixed data types in multiple sets of measurements.
problem Separating common and distinct information in mixed data types from multiple sources.
method Exponential Family Simultaneous Component Analysis (ESCA) model with structured sparse loading matrix.
result The proposed method effectively disentangles global, local common and distinct information.
Paper analyzes convergence of SGD with shuffling in distributed learning.
problem Analyzing convergence of distributed SGD with shuffling.
method Formalizes data partition with global/local shuffling, proves convergence for convex and non-convex cases, and considers insufficient shuffling.
result SGD with global shuffling converges in both convex and non-convex cases, and local shuffling is slower.
A new loss function α-loss bridges log-loss and 0-1 loss for binary classification.
problem Improving binary classification performance using a tunable loss function.
method Introducing α-loss, proving its margin-based form and classification-calibration, and providing an upper bound on empirical risk. result Empirical and expected risk difference upper bound for logistic regression-based classification.
Introduces Fitzpatrick losses, tighter than Fenchel-Young losses.
problem Improving loss functions for machine learning.
method Introduces Fitzpatrick losses based on the Fitzpatrick function.
result Fitzpatrick losses are tighter than Fenchel-Young losses.
We study losses for binary classification and class probability estimation and extend the understanding of them from margin losses to general composite losses which are the composition of a proper loss with a link function. We characterise when margin losses can be proper composite losses, explicitly show how to determ…
Tamed Cross Entropy (TCE) loss outperforms standard CE loss in noisy classification tasks.
problem Improving classification performance in noisy data scenarios.
method Introducing Tamed Cross Entropy (TCE) loss, a derivative of Cross Entropy (CE) loss.
result TCE loss outperforms CE loss in all tested noisy classification scenarios.
Unified surrogate loss framework for multi-label learning with strong consistency guarantees.
problem Improving consistency and accounting for label correlations in multi-label learning.
method Introducing multi-label logistic loss and extending it to comprehensive multi-label comp-sum losses, proving strong consistency guarantees for any multi-label loss.
result Unified surrogate loss framework benefiting from strong consistency guarantees for any multi-label loss.
This paper introduces new loss functions for balanced multi-class classification.
problem Balancing class imbalance in multi-class classification.
method Introduces two new surrogate loss families: GLA and GCA.
result GCA losses offer stronger theoretical guarantees in imbalanced settings.