KD can lead to student-teacher deviations that improve performance.
problem KD can lead to student-teacher deviations that may outperform the teacher.
method Characterized and explained the nature of student-teacher deviations through experiments and theory.
result KD can lead to improved generalization by exaggerating the implicit bias of gradient descent.
Proposes RaT to mitigate bias in student-teacher estimation.
problem Systematic bias in teacher's predictions propagates to student model.
method Uses teacher to estimate residuals in student's predictions.
result RaT method reduces teacher bias effect and achieves optimal rate.
Self-training with noisy student-teacher boosts keyword spotting accuracy.
problem Robust keyword spotting in challenging conditions.
method Aggressive data augmentation and self-training with noisy student-teacher approach.
result Significant accuracy improvement in difficult conditions, up to 60%.
Student-teacher learning improves generalization with noisy inputs.
problem Transfer knowledge from clean inputs to noisy inputs.
method Analyzes student-teacher learning using deep linear networks and experiments with nonlinear networks.
result Three factors are vital for success: zero training loss, teacher knowledge, and feature decomposition.
A new method detects concept drift without true labels.
problem Detecting concept drift in unsupervised settings.
method Student-teacher learning paradigm for drift detection.
result The method outperforms state-of-the-art approaches in experiments.
This work proposes a student-teacher network for predicting hospital admission locations.
problem Accurate prediction of hospital admission locations to optimize resource allocation.
method Reinforcement learning approach where a teacher network selects data batches for a student network.
result The approach outperforms state-of-the-art methods on tabular data and image recognition.
A new gradient flow for MMD with closed-form implementation.
problem Existing gradient flows either lack tractable numerical implementation or require strong assumptions.
method Introduces a (de)-regularized Maximum Mean Discrepancy (DrMMD) and its gradient flow.
result Guarantees near-global convergence for a broad class of targets in both continuous and discrete time.
Study shows RFRR's effectiveness with nearly orthogonal data in overparameterized settings.
problem Understanding the effectiveness of random feature regression with nearly orthogonal data.
method Investigates RFRR with nearly orthogonal deterministic unit-length input data vectors in the overparameterized regime.
result Shows high-probability non-asymptotic concentration results for RFRR's training, cross-validation, and generalization errors.
Framework models supervised learning in non-stationary data.
problem Non-stationary data in supervised learning.
method Statistical physics methods applied to LVQ and neural networks.
result LVQ and ReLU have different sensitivity to concept drift.
LoRA fine-tuning explained with gradient dynamics for low-rank perturbations.
problem Understanding why gradient descent converges to useful low-rank perturbations in LoRA fine-tuning.
method Generalized student-teacher setting with i.i.d. samples and online gradient descent.
result Gradient descent converges to the teacher model in dkO(1) iterations under certain conditions. We analyze the dynamics of an algorithm for approximate inference with large Gaussian latent variable models in a student-teacher scenario. To model nontrivial dependencies between the latent variables, we assume random covariance matrices drawn from rotation invariant ensembles. For the case of perfect data-model matc…
Introduces Star-Shaped deviation measures for risk analysis.
problem Risk measurement and analysis in finance.
method Characterizes Star-Shaped deviation measures through acceptance sets and convex deviation measures.
result Exposes the relationship between Star-Shaped risk measures and deviation measures.
Paper characterizes monotonic mean-deviation risk measures.
problem Developing consistent risk measures from mean-deviation models.
method Applying a risk-weighting function to the deviation part of a mean-deviation model.
result Characterizes monotonic mean-deviation measures as consistent risk measures.
We extend previous large deviations results for the randomised Heston model to the case of moderate deviations. The proofs involve the Gärtner-Ellis theorem and sharp large deviations tools.
Paper proves large deviation principle for stochastic approximations.
problem Asymptotic estimates of learning algorithm deviations.
method Weak convergence approach to large deviations.
result Identifies appropriate scaling sequence and new representation for rate function.
In this paper we propose the notion of dynamic deviation measure, as a dynamic time-consistent extension of the (static) notion of deviation measure. To achieve time-consistency we require that a dynamic deviation measures satisfies a generalised conditional variance formula. We show that, under a domination condition,…
Study large deviations in life insurance portfolios without identical distributions.
problem Large deviations in life insurance portfolios with bounded losses and variances.
method Upper bound from standard large deviations, counterexample for full large deviation principle.
result Exponential bound for average loss exceeding a threshold.
Study large deviations for hypoelliptic diffusion on sub-Riemannian manifolds.
problem Large deviations for hypoelliptic diffusion measures on sub-Riemannian manifolds.
method Rough path theory and manifold-valued Malliavin calculus.
result Proved a large deviation principle for pinned hypoelliptic diffusion measures.
Proposes new deviation measures using Minkowski gauges.
problem Lack of suitable acceptance sets for deviation measures.
method Derives deviation measures through Minkowski gauges of acceptable sets.
result Any positive homogeneous deviation measure can be accommodated in the framework.
In this paper we analyze a dynamic recursive extension of the (static) notion of a deviation measure and its properties. We study distribution invariant deviation measures and show that the only dynamic deviation measure which is law invariant and recursive is the variance. We also solve the problem of optimal risk-sha…
We provide a unifying treatment of pathwise moderate deviations for models commonly used in financial applications, and for related integrated functionals. Suitable scaling allows us to transfer these results into small-time, large-time and tail asymptotics for diffusions, as well as for option prices and realised vari…
Importance sampling has become an important tool for the computation of tail-based risk measures. Since such quantities are often determined mainly by rare events standard Monte Carlo can be inefficient and importance sampling provides a way to speed up computations. This paper considers moderate deviations for the wei…
Connections between Lie derivatives and the deviation equation has been investigated in spaces with affine connection. The deviation equations of the geodesics as well as deviation equations of non-geodesics trajectories have been obtained on this base. This is done via imposing certain conditions on the Lie derivative…
Unified approach to stochastic Volterra systems' deviations.
problem Large and moderate deviations for stochastic Volterra systems.
method Weak convergence approach by Budhijara, Dupuis and Ellis.
result Unified treatment of deviations for a broad class of stochastic Volterra equations.
Study large deviations in random walks on Lie groups.
problem Large deviations in sub-Riemannian random walks.
method Prove large deviation principle for random walks on stratified Lie groups.
result Proved a large deviation principle with a rate function adapted to sub-Riemannian geometry.
This is a report of our lessons learned building acoustic models from 1 Million hours of unlabeled speech, while labeled speech is restricted to 7,000 hours. We employ student/teacher training on unlabeled data, helping scale out target generation in comparison to confidence model based methods, which require a decoder…
Deviation inequalities and limit laws for random walks on metric spaces.
problem Understanding random walks on metric spaces with contracting isometries.
method Adapting Gouëzel's pivotal time construction to establish deviation inequalities.
result Exponential bounds and limit laws for random walks on mapping class groups and CAT(0) spaces.
Let M be a smooth manifold and S a semi-spray defined on a sub-bundle C of the tangent bundle TM. In this work it is proved that the only non-trivial k-jet approximation to the exact geodesic deviation equation of S, linear on the deviation functions and invariant under an spec…
Optimizes variance reduction in Heston model using large and moderate deviations.
problem Improving variance reduction in stochastic volatility models.
method Large and moderate deviations theory applied to Heston model.
result Derives closed-form solutions for optimal change of measure.
Large deviations theory applied to policy gradient methods.
problem Understanding convergence of policy gradient methods in reinforcement learning.
method Large deviation rate function and contraction principle from large deviations theory.
result Convergence properties of policy gradient methods can be extended to various policy parametrizations.
Study examines large deviations in random walks on hyperbolic spaces.
problem Large deviations in random walks on Gromov-hyperbolic spaces.
method Established large deviations results for distance and translation length of random walks.
result Deduced a special case of a conjecture regarding spectral radii of random matrix products.
Study large deviations and speed of random walks in hyperbolic spaces.
problem Understanding the speed of random walks in hyperbolic spaces.
method Large deviations analysis for random walks with a non-elementary semi-group.
result Established large deviations results for random walk distances.
Deviation inequalities for stochastic approximation methods.
problem Establishing bounds on the deviation of stochastic approximation methods.
method Martingale approximation method for separately Lipschitz functions.
result Established various deviation inequalities for stochastic approximation by averaging and minimization.
Large deviations for fat tailed distributions, i.e. those that decay slower than exponential, are not only relatively likely, but they also occur in a rather peculiar way where a finite fraction of the whole sample deviation is concentrated on a single variable. The regime of large deviations is separated from the regi…
We study layered neural networks of rectified linear units (ReLU) in a modelling framework for stochastic training processes. The comparison with sigmoidal activation functions is in the center of interest. We compute typical learning curves for shallow networks with K hidden units in matching student teacher scenarios…
We study a rolling model from the perspective of probability. More precisely, we consider a Riemannian manifold rolling against Euclidean space, where the rolling is coupled with random slipping and twisting. The system is modelled by a stochastic differential equation of Stratonovich-type driven by semimartingales, on…
Researchers introduce a method to assess the safety of interpretable machine learning models.
problem Ensuring safety in machine learning models that are easy to understand.
method Introduce maximum deviation as an optimization problem to find the largest deviation from a safe reference model.
result Interpretability helps in assessing the safety of machine learning models.
The displacement and deviation vectors in spaces (manifolds), the tangent bundle of which is endowed with a transport along paths, are introduced. In case these spaces are equipped with a linear connection, the deviation equations (between arbitrary, geodesic or not, paths) in such spaces are investigated.
Large deviation principle for deep neural networks with ReLU activation.
problem Understanding the behavior of deep neural networks with ReLU activation.
method Proving a large deviation principle for networks with Gaussian weights and ReLU activation functions.
result Simplified expressions and power-series expansions for the ReLU case.
Study large deviations in fractional volatility models with non-Gaussian volatility.
problem Large deviations in fractional volatility models with non-Gaussian volatility.
method Established a small-noise large deviation principle for log-price.
result Logarithmic call price asymptotics for large strikes in a special case.
In these notes, we present some methods and applications of large deviations to finance and insurance. We begin with the classical ruin problem related to the Cramer's theorem and give en extension to an insurance model with investment in stock market. We then describe how large deviation approximation and importance s…
The paper explores optimal insurance contracts using various deviation measures.
problem Optimal insurance contracts with mean-deviation measures.
method Study of convex signed Choquet integrals and standard deviation as deviation measures, analyzing premium principles like expected value, Value-at-Risk, and Expected Shortfall.
result Characterization of optimal indemnities and deductibles under different premium principles.
We establish large deviation principles for convolutional neural networks.
problem Understanding the behavior of convolutional neural networks in the infinite-channel limit.
method We establish large deviation principles for convolutional neural networks under Gaussian prior and posterior distributions.
result We provide a large deviation principle for the sequence of conditional covariance matrices and the posterior distribution.
Paper proposes a new daily benchmark for post-GFC government bond CIP deviations.
problem Lack of a canonical daily benchmark for CIP deviations.
method Used G10 plus KRW currency-tenor panels to analyze three lagged public state variables.
result Three lagged public state variables deliver strong performance in daily regressions.
The paper establishes a connection between different risk measures and their risk contributions.
problem Understanding the relationship between conditional coherent and deviation risk measures.
method Axiomatic framework and continuous-time risk contribution analysis.
result Risk contributions of time-consistent risk measures are also time-consistent.
Recently, it has been shown that Absolute Parallelism (AP) geometry admits paths that are naturally quantized. These paths have been used to describe the motion of spinning particles in a background gravitational field. In case of a weak static gravitational field limits, the paths are applied successfully to interpret…
The paper analyzes how a known density function can be deviated by a mixture distribution as more data is collected.
problem Modeling the deviation of a known density function when more data is collected.
method A novel distinguishability notion is used to establish rates of convergence for maximum likelihood estimates of the deviated proportion and latent mixing measure.
result Rates of convergence for the maximum likelihood estimates of the deviated proportion and latent mixing measure are established under the Wasserstein metric.
Large deviation principles for multivariate stochastic volatility models.
problem Understanding the behavior of log-processes in multivariate stochastic volatility models.
method Establishing a comprehensive sample path large deviation principle for log-processes.
result Asymptotic formulas for first exit times and barrier option prices derived from the LDP.