Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

6.3%12.5%18.8%25.0% · Apr 199419922001200920172026
48 results for prediction divergence

We extend a variational framework to estimate calibration errors for LpL_p divergences.

problem Ensuring predicted probabilities match observed class frequencies in machine learning.
method Extend variational framework to LpL_p divergences, separating over- and under-confidence.
result Avoids overestimation and separates over- and under-confidence.

Diverging Flows detects extrapolations in flow models, ensuring reliable predictions.

problem Flow models extrapolate into invalid data, leading to silent failures.
method Structurally enforce inefficient transport for off-manifold inputs.
result Effective detection of extrapolations without compromising predictive fidelity or inference latency.

Boosting framework for vector-valued prediction with geometric stability.

problem Lack of a general theoretical understanding of aggregation for structured prediction.
method Identifies (α,β)(α,β)-stability property and proposes a boosting framework based on exponential reweighting and geometric-median aggregation.
result Obtains exponential decay of empirical divergence error under weak learner condition and (α,β)(α,β)-stability.

The paper analyzes the bias-variance tradeoff for Bregman divergences.

problem Understanding the bias-variance tradeoff for Bregman divergences.
method Analyzes the bias-variance tradeoff through operations in dual space.
result Derives several results including a generalized law of total variance and ensembling operations.

Study optimizes tree-based models for better alignment of predicted scores and actual probabilities.

problem Traditional calibration metrics fail to align predicted scores with actual probabilities when score distributions deviate from the underlying data.
method Optimizes tree-based models (Random Forest, XGBoost) using Kullback-Leibler (KL) divergence to minimize the difference between predicted and true probability distributions.
result Optimized tree-based models yield superior alignment between predicted scores and actual probabilities without significant performance loss.

This paper analyzes privacy risks in neural network pruning and proposes a defense mechanism.

problem Privacy risks in neural network pruning due to membership inference attacks.
method Investigates the impact of pruning on prediction divergence and proposes a self-attention membership inference attack.
result Proposed defense mechanism mitigates privacy risks while maintaining sparsity and accuracy.

The Kelly Criterion is applied to prediction markets to analyze risk and return.

problem Mean beliefs in prediction markets often differ from actual prices.
method Logarithmic utility and Kullback-Leibler divergence are used to study risk and return adjustments.
result Misjudgment of bias and investment fraction affect portfolio growth rate.

We present a method for explaining the image classification predictions of deep convolution neural networks, by highlighting the pixels in the image which influence the final class prediction. Our method requires the identification of a heuristic method to select parameters hypothesized to be most relevant in this pred…

2017-11-17abs ↗pdf ↗

This work explores function-space inference using KL divergence and proposes Bayesian linear regression as a benchmark.

problem Approximating the predictive posterior distribution of Bayesian models without parameter posterior approximation.
method Employing Kullback-Leibler divergence and proposing featurized Bayesian linear regression as a benchmark.
result Minimizing KL divergence leads to an ill-defined objective function, highlighting limitations of this approach.

Estimates KL divergence with fairness considerations for sub-populations.

problem Fairly estimate KL divergence between distributions considering sub-populations.
method Proposes multi-group attribution for KL divergence estimation, derived from multi-calibration.
result Shows multi-group attribution provides better KL divergence estimates conditioned on sub-populations.

For nonlinear supervised learning models, assessing the importance of predictor variables or their interactions is not straightforward because it can vary in the domain of the variables. Importance can be assessed locally with sensitivity analysis using general methods that rely on the model's predictions or their deri…

2019-10-17abs ↗pdf ↗

Paper connects rejection learning to Bhattacharyya divergence.

problem Learning models to abstain from predictions.
method Developed a link between rejection and thresholding different statistical divergences, focusing on Bhattacharyya divergence.
result Rejector obtained by joint ideal distribution corresponds to thresholding of skewed Bhattacharyya divergence.

The problem of estimating an unknown discrete distribution from its samples is a fundamental tenet of statistical learning. Over the past decade, it attracted significant research effort and has been solved for a variety of divergence measures. Surprisingly, an equally important problem, estimating an unknown Markov ch…

2018-10-28abs ↗pdf ↗

StAD predicts divergence of diffusion and flow models without Jacobian computation.

problem Computing likelihood from diffusion and flow models is computationally expensive.
method Introduces StAD, a distillation method to predict divergence using Langevin-Stein operator.
result StAD predicts divergence with competitive variance and speed compared to existing methods.

We compare two models of corporate default by calculating the Jeffreys-Kullback-Leibler divergence between their predicted default probabilities when asset correlations are either high or low. Our main results show that the divergence between the two models increases in highly correlated, volatile, and large markets, b…

2016-04-24abs ↗pdf ↗

The study assesses low-rank approximations in Gaussian Process regression.

problem Improving the efficiency of Gaussian Process regression while maintaining accuracy.
method Analyzes two low-rank approximations: random Fourier features and Mercer expansion truncation, and bounds the divergence and error between exact and approximate models.
result Theoretical bounds on the divergence and error between exact and approximate Gaussian Process models are provided.

Generalizes bias-variance decomposition for Bregman divergences.

problem No specific problem stated; generalization of bias-variance for Bregman divergences.
method Provided a generalization of the bias-variance decomposition for Bregman divergences.
result A clear, standalone derivation of the bias-variance decomposition for Bregman divergences.

Study predicts price predictability in ultra-high frequency financial data using entropy tests.

problem Tackles predictability of ultra-high frequency financial data.
method Develops statistical tests based on Shannon entropy and Kullback-Leibler divergence to analyze predictability.
result Degree of randomness increases with aggregation level in transaction time.

We develop a method to combine Markov chain Monte Carlo (MCMC) and variational inference (VI), leveraging the advantages of both inference approaches. Specifically, we improve the variational distribution by running a few MCMC steps. To make inference tractable, we introduce the variational contrastive divergence (VCD)…

2019-05-10abs ↗pdf ↗

New loss functions based on f-divergences improve language model performance.

problem Improving multiclass classification and language modeling performance.
method Constructing new convex loss functions using f-divergences and deriving an operator for computation.
result The αα-divergence loss function with α=1.5α=1.5 performs well across various tasks.

New approach mitigates feedback divergence in imitation learning.

problem Divergence between held-out error and learner performance in imitation learning.
method Identifies covariate shift as the root cause and proposes a simulator-based solution.
result Naive behavioral cloning performs well in real-world decision making problems.

Mixability of a loss is known to characterise when constant regret bounds are achievable in games of prediction with expert advice through the use of Vovk's aggregating algorithm. We provide a new interpretation of mixability via convex analysis that highlights the role of the Kullback-Leibler divergence in its definit…

2014-03-10abs ↗pdf ↗

There has been a growing interest in mutual information measures due to their wide range of applications in Machine Learning and Computer Vision. In this paper, we present a generalized structured regression framework based on Shama-Mittal divergence, a relative entropy measure, which is introduced to the Machine Learn…

2014-09-26abs ↗pdf ↗

The study assesses low-rank approximations in Gaussian Process regression.

problem Improving Gaussian Process regression efficiency with low-rank approximations.
method Analyzes two low-rank approximations: random Fourier features and Mercer expansion truncation.
result Bounds on the divergence and error between exact and approximate GP models.

Estimates proper calibration errors and refinement terms in probabilistic predictions.

problem Lack of a general estimator for proper calibration errors and refinement terms with known statistical properties.
method Proposes a method for consistent, asymptotically unbiased estimation of proper calibration errors and refinement terms.
result Proves the relation between refinement and f-divergences, implying information monotonicity in neural networks.

Proposes a scalable method for counterfactual prediction using machine learning.

problem De-bias causal estimators with high-dimensional data in observational studies.
method Uses entropy balancing to learn weights minimizing Jensen-Shannon divergence, leading to robust counterfactual predictions.
result Consistent causal estimation if either propensity score or outcome model is correctly specified.

A new method learns posterior and predictive distributions together, reducing computational cost.

problem Sequential two-stage Bayesian inference is computationally expensive.
method Amortized variational inference targeting posterior-predictive distribution.
result Efficient online inference with more accurate predictive distributions.

CAKD framework optimizes knowledge transfer by focusing on influential components of distillation.

problem Balancing and optimizing knowledge transfer in distillation models.
method Decouple KL divergence into BCD, SCD, and WCD; prioritize influential components.
result CAKD framework consistently outperforms baseline across diverse models and datasets.

New method improves deep neural network performance in regression tasks.

problem Improving generalization, robustness, and explainability of deep neural networks in regression.
method Developed a new Information Bottleneck approach using Cauchy-Schwarz divergence.
result Demonstrated superior performance on six real-world regression tasks.

Theoretical proof shows COMs are a type of contrastive divergence model with improved sampling.

problem Improving sampling quality in offline model-based optimization.
method Showed COMs are contrastive divergence models, proposed Langevin MCMC sampler, and decoupled model.
result Improved sampling quality achieved by decoupling model and using Langevin MCMC.

New robust learning framework for regression NNs using β-divergences.

problem Outliers and data contamination in regression NNs training.
method Proposes rRNet based on β-divergence for robust learning of regression NNs.
result rRNet achieves optimal 50% asymptotic breakdown point for all β ∈ (0, 1].

Different optimizer choices lead to different financial model predictions.

problem The impact of optimizer choice on neural network models in financial time series.
method Analysis of large-scale volatility forecasting for S&P 500 stocks using various model-training-pipeline pairs.
result Optimizer choice reshapes non-linear response profiles and temporal dependence in financial models, leading to different functional outcomes.

This paper introduces a novel approach for learning to rank (LETOR) based on the notion of monotone retargeting. It involves minimizing a divergence between all monotonic increasing transformations of the training scores and a parameterized prediction function. The minimization is both over the transformations as well …

2012-10-16abs ↗pdf ↗

Reassesses calibration metrics in machine learning models.

problem Inconsistent reporting of calibration metrics in recent literature.
method Calibration-based decomposition of Bregman divergences, visualization of calibration and generalization error.
result New visualization technique for detecting trade-offs between calibration and generalization.

Black-box alpha (BB-αα) is a new approximate inference method based on the minimization of αα-divergences. BB-αα scales to large datasets because it can be implemented using stochastic gradient descent. BB-αα can be applied to complex probabilistic models with little effort since it only requires as input the likel…

2015-11-10abs ↗pdf ↗

Matrix SMD converges to unique solution minimizing Bregman divergence.

problem High-dimensional multi-output classification and matrix completion problems.
method Stochastic Mirror Descent with matrix parameters and matrix mirror functions.
result Matrix SMD converges exponentially to the unique solution minimizing Bregman divergence.