Rate GENERIC extends thermodynamics principles to non-equilibrium systems.
problem Understanding non-equilibrium thermodynamics and its relation to equilibrium thermodynamics.
method Developed a geometrical framework for rate GENERIC, extending Onsager's variational principle.
result Rate GENERIC structure provides a new perspective on thermodynamics in non-equilibrium systems.
SALR improves deep learning generalization by dynamically adjusting learning rates.
problem Improving generalization in deep learning models.
method Sharpness-aware learning rate scheduling based on local loss function sharpness.
result SALR drives solutions to flatter regions, improving generalization and convergence.
Learning rate annealing helps even in convex problems, improving generalization.
problem Improving generalization in neural networks, especially convex problems.
method Learning rate annealing schedule (large initial, then small learning rate).
result Gradient descent can reach minima with better generalization using learning rate annealing.
Stochastic gradient descent with a large initial learning rate is widely used for training modern neural net architectures. Although a small initial learning rate allows for faster training and better test performance initially, the large learning rate achieves better generalization soon after the learning rate is anne…
A novel approach models rating transitions using Lie groups and Deep Learning.
problem Modeling rating transitions with geometric properties and stochastic processes.
method Introducing Itô-SDEs on Lie groups, using TimeGAN for calibration, and examining rating matrix properties.
result The geometric approach using Lie groups and Deep Learning generates a good fit for rating transitions.
Variational autoencoders optimize an objective that combines a reconstruction loss (the distortion) and a KL term (the rate). The rate is an upper bound on the mutual information, which is often interpreted as a regularizer that controls the degree of compression. We here examine whether inclusion of the rate also acts…
In this paper, we give a new sharp generalization bound of lp-MKL which is a generalized framework of multiple kernel learning (MKL) and imposes lp-mixed-norm regularization instead of l1-mixed-norm regularization. We utilize localization techniques to obtain the sharp learning rate. The bound is characterized by the d…
This paper models short rates with jumps using PDEs.
problem Capturing jumps and spikes in interest rates.
method PDE approach for pricing interest rate derivatives.
result Established Feynman-Kač representation and derived solutions.
Large learning rates improve generalization, but optimal ranges are narrower than previously thought.
problem Optimizing learning rates for neural network training.
method Detailed exploration of learning rate ranges in a simplified setup, validating findings in a practical setting.
result Optimal learning rate ranges are significantly narrower than previously assumed.
Large learning rates improve neural network generalization, study shows.
problem Understanding why large learning rates lead to better neural network generalization.
method Visual analysis of training and testing loss landscapes, introduction of a nonlinear model.
result Extended phase with large learning rates leads to near-optimal generalization.
New bounds show polyhedral surrogates are optimal for generalization.
problem Proving generalization rates for polyhedral loss functions.
method Developed two general results for polyhedral surrogates.
result Polyhedral surrogates provide linear surrogate regret bounds, translating directly to target rates.
This study analyzes communication constraints in MoE architectures using information theory.
problem Communication constraints in Mixture-of-Experts (MoE) architectures.
method Developed a rate-distortion characterization of finite-rate gating in MoE architectures using information theory.
result Yielded capacity-aware limits for communication-constrained MoE systems.
We extend Dupire's formula for stochastic interest rates and local volatility.
problem Deriving formulas for stochastic interest rates and local volatility.
method Generalizations of Dupire's formula for stochastic drift and local volatility.
result Validated the limits of the generalized Dupire formulae for specific cases.
Study on optimal rates for sequential probability assignment using smoothed analysis.
problem Optimal rates for sequential probability assignment under smoothed adversaries.
method General-purpose reduction from minimax rates to transductive learning, development of an efficient algorithm using MLE oracle.
result Optimal (logarithmic) fast rates for parametric and finite VC dimension classes, sublinear regret for general classes.
New models for short rates show longer periods at higher rates.
problem Modeling longer periods of higher interest rates.
method Developed a class of time-homogeneous one-factor Markov diffusion models with specific boundary conditions.
result Explicit expressions for bond prices and transition densities in new probability measure.
Large learning rates cause oscillations in NN weights that improve generalization.
problem Improving generalization of neural networks trained with large learning rates.
method Theoretical analysis and feature-noise data generation model.
result Oscillating SGD with large learning rates benefits NN generalization by effectively learning weak features.
Defines a new short rate model and convexity adjustment formulae.
problem Interest rate convexity in a Gaussian framework.
method Defines a short rate model driven by a Gaussian Volterra process and derives convexity adjustment formulae.
result Explicit formulae for convexity adjustment derived.
We present a general framework for solving a large class of learning problems with non-linear functions of classification rates. This includes problems where one wishes to optimize a non-decomposable performance metric such as the F-measure or G-mean, and constrained training problems where the classifier needs to sati…
The paper tackles fast rates in structured prediction problems.
problem Structured prediction problems with discrete outputs.
method Introducing continuous surrogate problems and leveraging their convergence rates for discrete problems.
result Super fast rates, including exponential rates, for excess risk in structured prediction problems.
New theorem for generalized group sparsity improves consistency and convergence rates.
problem Improving statistical inference in high-dimensional data with element-wise and group-wise sparsity.
method Developed a generalized version of Sparse-Group Lasso and proved a universal theorem for consistency and convergence rates.
result Obtained results on consistency and convergence rates for different forms of double sparsity regularization.
This article studies the achievable guarantees on the error rates of certain learning algorithms, with particular focus on refining logarithmic factors. Many of the results are based on a general technique for obtaining bounds on the error rates of sample-consistent classifiers with monotonic error regions, in the real…
Study problem-dependent rates in statistical learning theory, achieving optimal generalization error bounds.
problem Generalization error in statistical learning theory.
method Uniform localized convergence framework.
result Optimal generalization error bounds for various learning problems.
There are more than eight hundred interest rates published in China bond market every day. Which are the benchmark interest rates that have broad influences on most interest rates is a major concern for economists. In this paper, multi-variable Granger causality test is developed and applied to construct a directed net…
We study convergence rates of variational posterior distributions for nonparametric and high-dimensional inference. We formulate general conditions on prior, likelihood, and variational class that characterize the convergence rates. Under similar "prior mass and testing" conditions considered in the literature, the rat…
The study proposes algorithms to minimize rating discordance in missing data.
problem Missing ratings in combined rating lists.
method Optimization models and algorithms that minimize total rating discordance.
result The proposed methods outperform state-of-the-art imputation methods in accuracy.
This paper investigates the supervised learning problem with observations drawn from certain general stationary stochastic processes. Here by \emph{general}, we mean that many stationary stochastic processes can be included. We show that when the stochastic processes satisfy a generalized Bernstein-type inequality, a u…
Paper improves learning rates for SGD and NAG.
problem Generalization performance of stochastic optimization algorithms.
method Establishes new learning rates for SGD and NAG.
result Improved guarantees in some settings or comparable rates under weaker assumptions.
Improved SGD methods converge faster for nonconvex optimization.
problem Nonconvex optimization challenges in machine learning.
method Adaptive SGD with line-search and Polyak stepsizes.
result Unified convergence rates for various nonconvex functions.
Paper finds wide minima are better for generalization and proposes a new learning rate schedule.
problem The challenge of finding optimal learning rates for model training.
method The paper introduces a new hypothesis about the density of wide minima and designs an explore-exploit learning rate schedule.
result The explore-exploit learning rate schedule improves model performance and reduces training time.
We introduce here for the first time the long-term swap rate, characterised as the fair rate of an overnight indexed swap with infinitely many exchanges. Furthermore we analyse the relationship between the long-term swap rate, the long-term yield, see Biagini et al. [2018], Biagini and Härtel [2014], and El Karoui et a…
New insights into network generalization show learning rate affects both norm and sharpness.
problem Understanding the generalization of overparameterized networks.
method Empirical analysis and theoretical proof of the trade-off between norm and sharpness.
result Learning rate influences both norm and sharpness, neither alone minimizes generalization error.
The idea of forward rates stems from interest rate theory. It has natural connotations to transition rates in multi-state models. The generalization from the forward mortality rate in a survival model to multi-state models is non-trivial and several definitions have been proposed. We establish a theoretical framework f…
Study on convergence rates for optimal transport with regularization.
problem Convergence analysis of divergence-regularized optimal transport.
method Novel methodology using quantization and martingale couplings.
result Sharp rates for various divergences and transport costs.
Study optimal portfolio strategies with time-varying discount rates.
problem Optimizing portfolio decisions with a non-constant discount rate.
method Introduced subgame perfect strategies to handle time inconsistency, using fixed point iteration to find the utility-weighted discount rate.
result Subgame perfect strategies are equivalent to optimal strategies under certain utility function assumptions.
New learning rate approach reveals phase transitions in SGD performance.
problem Understanding feature learning dynamics in neural networks.
method Characterizing the relationship between learning rate(s) and sample complexity for gradient-based algorithms.
result Phase transition from information exponent to generative exponent regime with different learning rates.
Diffusion models improve image compression at low bit-rates.
problem Efficiently compressing images at very low bit-rates.
method Encoding into an embedding, using diffusion models to refine the embedding iteratively.
result Realistic reconstructions can be generated at extremely low bit-rates.
Neural models learn continuous-time Markov chain transition rates from data.
problem Learning transition rates for complex stochastic systems.
method Neural networks to model nonlinear transition rates from observed data.
result Neural models outperform traditional methods in accuracy.
Study affine models for alternative risk-free rates and derive caplet pricing formulas.
problem Valuation of caplets/floorlets in models for alternative risk-free rates.
method Affine process for RFRs, explicit valuation formulas for various derivatives.
result Explicit formulas for caplet/floorlet pricing in affine models for RFRs.
Bayesian method with Gaussian process priors achieves optimal convergence rates for regression function and its derivatives.
problem Estimating the regression function and its derivatives in nonparametric regression.
method Bayesian approach with Gaussian process priors, focusing on convergence rates and plug-in property.
result Equivalence of convergence rates of posterior distributions and Bayes estimators for regression function and its derivatives.
The well-known theorem of Dybvig, Ingersoll and Ross shows that the long zero-coupon rate can never fall. This result, which, although undoubtedly correct, has been regarded by many as surprising, stems from the implicit assumption that the long-term discount function has an exponential tail. We revisit the problem in …
Paper analyzes mistake and generalization of MNIC classifiers.
problem Understanding the performance of interpolating classifiers.
method Elementary analyses of MNIC's regret and generalization.
result MNIC generalizes with a rate proportional to the norm of the interpolating solution and inversely proportional to the number of data points.
Continuous, ubiquitous monitoring through wearable sensors has the potential to collect useful information about users' context. Heart rate is an important physiologic measure used in a wide variety of applications, such as fitness tracking and health monitoring. However, wearable sensors that monitor heart rate, such …
Paper analyzes faster convergence rates for reinforcement learning from offline data.
problem Analyzing faster convergence rates for reinforcement learning from offline data.
method Fine analysis of reinforcement learning from offline data, providing fast rates for regret convergence.
result The paper provides fast rates for the regret convergence, showing that the level of exponentiation depends on the noise in the decision-making problem.
We determine an explicit formula for the Laplace transform of the price of an option on a maximal interest rate when the instantaneous rate satisfies Cox-Ingersoll-Ross's model. This generalizes considerably one result of Leblanc-Scaillet.
The paper analyzes how learning rate affects SGD and provides insights into optimal rates.
problem Understanding the impact of learning rate on stochastic gradient descent.
method Developed a learning-rate-dependent stochastic differential equation (lr-dependent SDE) to analyze SGD.
result Established a linear rate of convergence for SGD and found the optimal linear rate by analyzing the spectrum of the Witten-Laplacian.
We present new excess risk bounds for general unbounded loss functions including log loss and squared loss, where the distribution of the losses may be heavy-tailed. The bounds hold for general estimators, but they are optimized when applied to η-generalized Bayesian, MDL, and empirical risk minimization estimators. …
Unified derivation of PAC-Bayes and MI bounds for general VC classes with fast rates.
problem Generalization bounds for machine learning models with VC classes.
method Unified derivation of conditional PAC-Bayesian and mutual information bounds, including MAC-Bayesian bounds.
result Nontrivial bounds for general VC classes and faster rates for specific conditions.
Model shows how discount rates affect intergenerational equity in climate mitigation.
problem Intergenerational equity in climate mitigation decisions.
method Extended DICE model with stochastic discount rates and financing extensions.
result Discount-rate uncertainty amplifies intergenerational inequality in climate mitigation.