Tilting loss functions improves machine learning performance.
problem Improving machine learning models, especially in under- and over-parameterized networks.
method Using evolving loss functions that emphasize different classes cyclically.
result Dynamical loss functions lead to better generalization and stability in training.
New loss function helps learn unstable dynamical systems.
problem Gradient descent fails to learn unstable dynamical systems.
method Introduced a time-weighted logarithmic loss function.
result Time-weighted loss function effectively learns unstable systems.
Teaches AI models to learn effectively through dynamic loss functions.
problem Optimizing machine learning models' performance through dynamic loss functions.
method Develops a method for a teacher model to dynamically output loss functions for a student model, enabling gradient-based optimization.
result Significantly improves the performance of various student models in real-world tasks.
SoftAdapt dynamically adjusts loss weights for multi-part functions.
problem Slow convergence and poor weight selection for multi-part loss functions.
method SoftAdapt dynamically changes weights based on live performance statistics.
result Improved convergence and better weight selection for multi-part loss functions.
New method learns stochastic thermodynamics from system currents.
problem Understanding entropy production in complex dynamical systems.
method Constructs learning framework using currents and machine learning loss functions.
result Derives loss functions for thermodynamic functions directly from dynamics.
New algorithm controls systems with unknown, changing losses.
problem Control systems with adversarial perturbations and unknown loss function.
method Efficient sublinear regret algorithm for bandit convex optimization with memory.
result Achieves efficient control with sublinear regret in the presence of unknown, changing losses.
Reduces dynamic regret to static problem in RKHS.
problem Minimizing cumulative loss in online convex optimization.
method Reduces dynamic regret to static regret problem in RKHS.
result Optimal dynamic regret guarantees for linear losses and new bounds for exp-concave and improper linear regression.
This study addresses the challenges of dynamic mini-batch sub-sampling in neural network training.
problem Challenges in training neural networks due to dynamic mini-batch sub-sampling.
method Distinguishes between static and dynamic sub-sampling, recasting optimization to find SNN-GPPs.
result SNN-GPPs are less susceptible to sub-sampling-induced discontinuities and better approximate true optima.
Optimal algorithms for mixable losses in dynamic environments with reduced redundancy.
problem Online optimization of mixable loss functions in a dynamic environment.
method Introduce online mixture schemes with polynomial and logarithmic time complexities.
result Achieves optimal redundancy up to a constant multiplicity gap.
New algorithms reduce dynamic regret for convex and smooth functions in non-stationary environments.
problem Online convex optimization in non-stationary environments.
method Proposed novel online algorithms exploiting smoothness to reduce dynamic regret.
result Dynamic regret improved to O ( T ) \mathcal{O}(T) O ( T ) for convex and smooth functions. Transformers struggle to learn Markovian dynamics, showing NP-hard optimization challenges.
problem Understanding transformers' limitations in learning Markovian dynamical functions.
method Investigated through a structured ICL setup, analyzing loss landscapes and parameter optimization.
result Recovering optimal transformer parameters for Markovian functions is NP-hard.
New method measures model risk in dynamic settings with uncertain state processes.
problem Lack of non-parametric approach for dynamic model risk quantification.
method Generalizes relative-entropic approach to dynamic case under f f f -divergence. result Unified treatment for worst-case risk and f f f -divergence budget. This paper improves forecast stability without sacrificing accuracy using dynamic loss weighting.
problem Rolling origin forecast instability in time series forecasting.
method Dynamic loss weighting algorithms applied to the N-BEATS model.
result Dynamic loss weighting can further improve forecast stability without compromising accuracy.
Existing approaches to online convex optimization (OCO) make sequential one-slot-ahead decisions, which lead to (possibly adversarial) losses that drive subsequent decision iterates. Their performance is evaluated by the so-called regret that measures the difference of losses between the online solution and the best ye…
We discover scaling laws for kernel regression loss under various learning rate schedules.
problem Understanding loss dynamics and learning rate schedules in kernel regression.
method Theoretical analysis of stochastic gradient descent on a power-law kernel regression model.
result Established a Functional Scaling Law (FSL) capturing the full loss trajectory under arbitrary learning rate schedules.
We propose a simple discrete time semi-supervised graph embedding approach to link prediction in dynamic networks. The learned embedding reflects information from both the temporal and cross-sectional network structures, which is performed by defining the loss function as a weighted sum of the supervised loss from past…
Algorithm minimizes loss and constraint violations in online convex optimization with smooth penalties.
problem Minimizing loss and constraint violations in online convex optimization with smooth penalties.
method Projected gradient descent over a set around the current action.
result Both dynamic regret and constraint violation are bounded by the path-length.
Inverse depth scaling found in LLMs due to similar layers averaging error.
problem Understanding how depth affects loss in large language models.
method Analysis of LLMs and toy residual networks.
result Loss scales inversely proportional to depth in LLMs.
The paper introduces a multi-step loss function to improve model-based reinforcement learning.
problem Compounding of one-step prediction errors in long trajectories.
method A multi-step objective function combining MSE losses at various future horizons.
result Models trained with the multi-step loss achieve significant improvement in future prediction.
This work focuses on dynamic regret of online convex optimization that compares the performance of online learning to a clairvoyant who knows the sequence of loss functions in advance and hence selects the minimizer of the loss function at each step. By assuming that the clairvoyant moves slowly (i.e., the minimizers c…
The paper optimizes daily storage trading of electricity using dynamic spread densities.
problem Optimizing daily storage trading of electricity based on price spreads.
method Formulated dynamic density functions based on skewed-t representations to model hourly electricity price spreads. Selected the best specification for each spread using the Pinball Loss function and calculated risk associated with spread arbitrages.
result Optimal daily operation of a battery storage facility determined from spread densities.
Adaptive loss function improves performance by aligning training and evaluation metrics.
problem Loss-metric mismatch in machine learning training.
method Adaptive loss alignment through meta-learning of a dynamic loss function.
result Significant performance improvements across various tasks and data.
Near-logarithmic regret per switch achieved for mixable/exp-concave losses.
problem Online optimization of mixable loss functions with dynamic environments.
method Online mixture framework using static solvers and hyper-expert creations.
result Near-logarithmic regret per switch with sub-polynomial complexity.
Study on privacy leakage in noisy gradient descent algorithms.
problem Information leakage of iterative randomized learning algorithms about training data.
method Analyzes the dynamics of Rényi differential privacy loss in noisy gradient descent algorithms.
result Privacy loss converges exponentially fast for smooth and strongly convex loss functions.
A new method automatically and dynamically sets learning rates in deep learning.
problem Determining the appropriate learning rate in deep learning tasks is challenging and often subjective.
method Local Quadratic Approximation (LQA) to automatically and dynamically set learning rates.
result The proposed method leads to nearly optimal learning rates in a computationally efficient way.
HydaLearn dynamically adjusts task weights for better MTL performance.
problem Constant loss weights in MTL lead to poor results due to drifting relevance and varying mini-batch composition.
method HydaLearn uses mini-batch gradients to dynamically adjust task weights.
result HydaLearn improves performance on synthetic and real-world data.
Study visualizes actor-critic loss landscapes for inventory optimization.
problem Difficulties in solving multi-store dynamic inventory control problems.
method Low-dimensional visualizations of actor loss function.
result Loss landscapes favor optimal policies in reinforcement learning.
The study improves VaR forecast accuracy by modeling conditional quantile dynamics.
problem Improving the accuracy of Value-at-Risk (VaR) forecasts for time-varying quantiles.
method Time-varying modeling of VaR, evaluation via simulation, asymmetric Mean Absolute Deviation loss function.
result Substantial improvements in forecasting conditional quantiles by maintaining predicted quantile unchanged.
Optimal liquidation strategy with dynamic risk adjustment in financial markets.
problem Maximizing profit and loss in risky security liquidation with price impact.
method Formulates as a stochastic optimal control problem, uses dynamic risk measures.
result Closed-form solution for optimal liquidation policies under quadratic driver.
We describe loss surfaces using topological Betti numbers.
problem Understanding the complexity and structure of loss surfaces in neural networks.
method Topological analysis using Betti numbers for multilayer neural networks.
result Loss complexity is influenced by the number of hidden units and activation function.
New insights into CE dynamics reveal how Hadamard initialization simplifies softmax.
problem Understanding the dynamics of cross-entropy training loss in deep learning.
method Analyzing a two-layer linear neural network with standard-basis vectors as inputs.
result Gradient flow on cross-entropy converges to neural collapse geometry, proving global convergence.
GOALS improves learning rate selection for dynamic MBSS in deep learning.
problem Challenges in selecting learning rates for dynamic MBSS in deep learning.
method Gradient-only approximation line search (GOALS) for dynamic MBSS loss functions.
result GOALS reduces model errors in multimodal cases.
This paper considers portfolio construction in a dynamic setting. We specify a loss function comprised of utility and complexity components with an unknown tradeoff parameter. We develop a novel regret-based criterion for selecting the tradeoff parameter to construct optimal sparse portfolios over time.
A new AI optimization method uses energy-conserving dynamics inspired by Born-Infeld theory.
problem Optimization challenges in non-convex loss functions and machine learning tasks.
method Discretization of Born-Infeld dynamics for energy-conserving Hamiltonian optimization.
result The method avoids high local minima and outperforms traditional methods in shallow valleys.
Deep learning dynamics exhibit anomalous superdiffusion initially, aiding escape from local minima.
problem Understanding the dynamics of learning in deep neural networks.
method Novel analysis of SGD dynamics and loss landscape structure.
result SGD exhibits anomalous superdiffusion initially, transitioning to subdiffusion as learning progresses.
Framework to generalize impermanent loss for decentralized exchanges.
problem Difficult analysis of impermanent loss due to diverse market maker algorithms and fee structures.
method Developed a framework to generalize impermanent loss for constant function market makers with optional concentrated liquidity.
result Identified conditions for profitability of liquidity provisioning.
SA algorithms control dynamic regret in non-stationary settings with strong convexity or exp-concavity.
problem Non-stationary Online Convex Optimization with dynamic regret control.
method Strongly Adaptive (SA) algorithms view dynamic regret as path variation of the comparator sequence.
result SA algorithms achieve i l d e O ( T V T ∨ log T ) ilde O(\sqrt{TV_T} \vee \log T) i l d e O ( T V T ∨ log T ) and i l d e O ( d T V T ∨ d log T ) ilde O(\sqrt{dTV_T} \vee d\log T) i l d e O ( d T V T ∨ d log T ) dynamic regret for strongly convex and exp-concave losses, respectively. Paper models and forecasts intra-day electricity price spreads.
problem Forecasting intra-day price spreads for electricity traders and operators.
method Dynamic density functions based on skewed-t distributions, conditional on exogenous drivers.
result Best fitting and forecasting specifications selected using Pinball Loss function.
Gradient descent optimizes neural networks and random features similarly, achieving zero loss fast.
problem Optimizing two-layer neural networks and random feature models under gradient descent.
method Comprehensive analysis of gradient descent dynamics, considering various network widths and data sizes.
result Gradient descent achieves zero training loss exponentially fast in the over-parametrized regime.
DILATE improves deep time series forecasting for non-stationary signals.
problem Forecasting non-stationary signals with multiple future steps.
method Introduces DILATE, a new objective function for deep neural nets.
result DILATE outperforms MSE and DTW in various non-stationary datasets.
TRS-ODENs learn dynamics with time-reversal symmetry for more efficient learning.
problem Learning dynamics with time-reversal symmetry for more efficient learning.
method Proposed a loss function and a new framework (TRS-ODENs) to learn dynamics efficiently.
result TRS-ODENs can learn dynamics from noisy and complex trajectories efficiently.
New setup for online learning captures continuous changes in losses, improving dynamic regret analysis.
problem Capturing regularity in online learning problems with continuous changes in losses.
method Introducing Continuous Online Learning (COL) and proving its equivalence to solving certain equilibrium problems (EPs).
result Achieving sublinear dynamic regret in COL is equivalent to solving certain EPs, offering conditions for efficient algorithms.
This paper addresses tracking of a moving target in a multi-agent network. The target follows a linear dynamics corrupted by an adversarial noise, i.e., the noise is not generated from a statistical distribution. The location of the target at each time induces a global time-varying loss function, and the global loss is…
Value functions struggle to represent transition dynamics, impacting statistical efficiency.
problem Limited representational power of value functions in capturing transition dynamics.
method Case studies of various reinforcement learning problems to explore the limitations of value-based methods.
result Value-based methods can be as efficient as model-based ones in some cases but severely underperform in others due to information loss.
Paper proposes self-supervised method for accurate speaker diarization.
problem Speaker diarization without massive labeling effort.
method Introduces dynamic triplet loss and multinomial loss for audio-video synchronization.
result Best model yields +8% F1-score improvement and diarization error rate reduction.
As deep Variational Auto-Encoder (VAE) frameworks become more widely used for modeling biomolecular simulation data, we emphasize the capability of the VAE architecture to concurrently maximize the timescale of the latent space while inferring a reduced coordinate, which assists in finding slow processes as according t…
DiMS sampler explores neural network loss minima via dissipative dynamics.
problem Sampling reparameterization invariant solutions in neural networks.
method Dynamical system based on kinetic energy with dissipative friction.
result DiMS sampler samples exactly from minimum level sets.
Enhances traffic forecasting with dynamic regression incorporating error modeling.
problem Improving accuracy of traffic forecasts using deep spatiotemporal models.
method Integrates matrix-variate autoregressive (AR) model into loss function for error series of base model.
result Improved traffic forecasting performance on SOTA models with interpretable AR coefficients.