A new method normalizes flow mixtures for better inference across different data types.
problem Inference failure across diverse posterior geometries in normalizing flows.
method Introduces a two-stage framework with a stable global weighting mechanism based on sEMA.
result Achieves consistent NLL improvements and stable weight trajectories over baselines.
Explains the difference between EMA and moving EMA, focusing on market trend indicators.
problem Understanding the difference between exponential moving average and moving exponential average.
method Explains the mathematical tools and definitions of trend indicators.
result Discusses the properties of the MACD indicator and its use in market trend analysis.
New moving average adapts weight dynamically based on polynomial and wavefunction.
problem Lagging traditional moving averages in adjusting to changes in data.
method Develops a moving average with weight as a polynomial of a wavefunction from an eigenproblem.
result Immediate 'switch' without lag, adapting to changes in data.
A new method for exponentially weighted moving models using approximations.
problem Efficiently updating moving averages for time series data.
method Approximates EWMM using a fixed window and quadratic term, solving non-growing problems.
result Approximation produces estimates similar to exact EWMM.
Paper proposes an online adaptation algorithm for improving model performance.
problem Improving model fidelity in real-time for domain shift and time variance.
method Extended Kalman Filter with Exponential Moving Average and Dynamic Multi-Epoch strategy.
result Proposed algorithm outperforms existing methods in experiments.
Improved averaging method for noisy observations converges strongly.
problem Noisy observations from random dynamical systems require stable estimates.
method Introduced p-EMA, a modified exponential moving average with subharmonic weight decay. result Stochastic convergence guarantees for p-EMA under mild assumptions. PACE optimizes training for averaged language models, improving performance.
problem How to optimize training for averaged language model iterates.
method Formulated as an optimal-control problem, solved for minimizing error of the average with a penalty on intervention size.
result PACE improves the limiting squared error of the iterate-average estimator by an arbitrarily large factor on some instances.
Paper analyzes how EMA improves SGD in linear regression.
problem Understanding the effectiveness of EMA in training deep learning models.
method Established risk bound for online SGD with EMA in linear regression.
result SGD with EMA has smaller variance error and exponentially decaying bias error.
We introduce a covariance matrix estimator that both takes into account the heteroskedasticity of financial returns (by using an exponentially weighted moving average) and reduces the effective dimensionality of the estimation (and hence measurement noise) via techniques borrowed from random matrix theory. We calculate…
We propose an explicit recursive method to approximate a power-law with a finite sum of weighted exponentials. Applications to moving averages with long memory are discussed in relationship with stochastic volatility models.
BEMA reduces bias in EMA, leading to faster convergence and better performance.
problem Stochasticity in language model fine-tuning destabilizes training.
method Bias-Corrected Exponential Moving Average (BEMA) augmentation of EMA.
result BEMA leads to significantly improved convergence rates and final performance.
Adaptive estimation for nonstationary time series reduces computational cost.
problem Estimating parameters of nonstationary time series with varying parameters over time.
method Moving exponential moving ML estimator for scale parameter estimation.
result Significantly improved log-likelihoods compared to standard estimation.
Machine learning models outperform traditional technical analysis in Bitcoin trading.
problem Maximizing profits in the Bitcoin market using trading signals.
method Comparison of machine learning models (LightGBM, LSTM) and technical analysis strategies (EMA, MACD+ADX).
result LSTM model achieved a 65.23% cumulative return over a year, significantly outperforming other strategies.
Moment polytope of toric exponential families is a projection of a simplex.
problem Understanding the geometry of exponential families in finite sample spaces.
method Toric torification and projection of higher-dimensional simplices.
result Moment polytope is a projection of a higher-dimensional simplex.
Adaptive t-distribution estimates nonstationary time series using moving moments.
problem Nonstationary time series with varying dependence structure.
method Moving estimator optimizing a weighted log-likelihood, using exponential moving averages for moments.
result Evolution of ν parameter in Student's t-distribution, capturing tail behavior and extreme events.
STORM-PG uses momentum for faster policy gradient updates.
problem Improving policy gradient methods for reinforcement learning.
method Introduces STORM-PG, a SARAH-based algorithm with exponential moving average.
result Achieves O(1/ε3) sample complexity, matching best-known rate. We examine two different techniques for parameter averaging in GAN training. Moving Average (MA) computes the time-average of parameters, whereas Exponential Moving Average (EMA) computes an exponentially discounted sum. Whilst MA is known to lead to convergence in bilinear settings, we provide the -- to our knowledge …
GPA improves LLM training speed by 8.71% for Llama-160M models.
problem Training Large Language Models (LLMs) with high memory overhead and slow convergence.
method Generalized Primal Averaging (GPA) extends Nesterov's method to eliminate memory-intensive two-loop structure.
result GPA achieves up to 10.13% speedup over AdamW in training Llama-1B model.
The paper triangulates Heisenberg groups with horizontal and straight simplexes.
problem Triangulating Heisenberg groups with specific regularity properties.
method Constructing triangulations with horizontal and straight simplexes on a polyhedral structure and extending to the whole Heisenberg group.
result Explicit examples of grid and triangulations provided.
Classifying streaming data requires the development of methods which are computationally efficient and able to cope with changes in the underlying distribution of the stream, a phenomenon known in the literature as concept drift. We propose a new method for detecting concept drift which uses an Exponentially Weighted M…
Improving optimization for iterate-averaged language models
problem How to optimize the averaged model returned by Language Model pipelines
method Formulating optimizer design as an optimal-control problem
result Proven convergence rate and strict improvement in squared error
Novel simplex-valued distribution improves on existing models.
problem Limitations of existing simplex-valued distributions like Dirichlet.
method Introducing continuous categorical distribution.
result Continuous categorical resolves limitations of Dirichlet.
We introduce a new distance metric for non-linear embeddings of Tempered Exponential Measures.
problem Non-linear embeddings of Tempered Exponential Measures (TEMs).
method Parameterization of finite discrete TEMs via Legendre functions, introducing tempered Hilbert co-simplex distance.
result Established a generalization of the Hilbert log cross-ratio simplex distance to a tempered Hilbert co-simplex distance.
The paper optimizes portfolios using MACD signals derived from price history.
problem Optimizing risky asset portfolios with latent mean-reverting and momentum factors.
method Derives optimal strategies based on MACD signals from EMA processes.
result Establishes admissibility and verification of optimal strategies.
Improved score-based models generate high-quality images up to 256x256.
problem Training score-based models for high-resolution images is unstable and limited.
method Theoretical analysis, exponential moving average of model weights.
result Score-based models can generate high-fidelity images up to 256x256.
New adaptive methods solve weakly convex stochastic optimization problems.
problem Solving weakly convex stochastic optimization problems.
method Adaptive first and zeroth-order methods using exponential moving averages.
result Established non-asymptotic convergence rates for nonsmooth and nonconvex problems.
Several recently proposed stochastic optimization methods that have been successfully used in training deep networks such as RMSProp, Adam, Adadelta, Nadam are based on using gradient updates scaled by square roots of exponential moving averages of squared past gradients. In many applications, e.g. learning with large …
We build a multiassets heterogeneous agents model with fundamentalists and chartists, who make investment decisions by maximizing the constant relative risk aversion utility function. We verify that the model can reproduce the main stylized facts in real markets, such as fat-tailed return distribution and long-term mem…
We study capital process behavior in the fair-coin game and biased-coin games in the framework of the game-theoretic probability of Shafer and Vovk (2001). We show that if Skeptic uses a Bayesian strategy with a beta prior, the capital process is lucidly expressed in terms of the past average of Reality's moves. From t…
We show Vector Autoregressive Moving Average models with scalar Moving Average components could be estimated by generalized least square (GLS) for each fixed moving average polynomial. The conditional variance of the GLS model is the concentrated covariant matrix of the moving average process. Under GLS the likelihood …
A geometric triangulation of a Riemannian manifold is a triangulation where the interior of each simplex is totally geodesic. Bistellar moves are local changes to the triangulation which are higher dimensional versions of the flip operation of triangulations in a plane. We show that geometric triangulations of a compac…
We propose that a simple, Lagrangian 2d N=(0,2) duality interface between the 3d N=2 XYZ model and 3d N=2 SQED can be associated to the simplest triangulated 4-manifold: the 4-simplex. We then begin to flesh out a dictionary between more general triangulated 4-manifolds with boundar…
A new method Expectigrad improves on Adam and RMSProp by reducing divergence and improving performance.
problem Improving the convergence properties of adaptive gradient methods like Adam and RMSProp.
method Adjusts stepsizes using a per-component unweighted mean of all historical gradients and a bias-corrected momentum term.
result Cannot diverge on convex optimization problems that cause Adam to diverge.
This paper provides an insight to the time-varying dynamics of the shape of the distribution of financial return series by proposing an exponential weighted moving average model that jointly estimates volatility, skewness and kurtosis over time using a modified form of the Gram-Charlier density in which skewness and ku…
Behavior cloning training instabilities amplified by SGD noise over long horizons.
problem Training instabilities in behavior cloning with deep neural networks.
method Empirical dissection of minibatch SGD updates and their effects on long-horizon rewards.
result Exponential moving average (EMA) of iterates effectively mitigates gradient variance amplification (GVA).
A new optimization method for probability simplex problems.
problem Optimizing convex problems over the probability simplex.
method Cauchy-Simplex iteration scheme, mapping to sphere, gradient descent, and back-mapping.
result Convergence results and faster convergence in high dimensions.
Time series analysis is a key component of machine learning, with applications in various fields.
problem Time series analysis in machine learning
method Basic concepts, classical statistical models, modern machine learning approaches
result Machine learning techniques for time series analysis
Bayesian model predicts evolving guest origin markets in tourism.
problem Forecasting the changing composition of guest origin markets in tourism.
method Developed and applied Bayesian Dirichlet autoregressive moving average (BDARMA) models to Airbnb booking data.
result BDARMA models achieve lower forecast error and competitive performance in guest origin market shares.
Auto-regressive models improve smoothing efficiency with exponentially tapered windows.
problem Improving time-series smoothing efficiency.
method An auto-regressive formulation for time-series smoothing.
result Auto-regressive models result in moving means with exponentially tapered windows.
This study uses moving average cluster entropy to analyze financial market dynamics.
problem Understanding long-range dependence in financial markets.
method Moving average cluster entropy approach applied to ARFIMA and FBM processes.
result Long-range positive correlation in financial markets is linked to the cluster entropy behavior.
Bayesian models predict evolving guest origin markets in tourism.
problem Forecasting the changing composition of guest origin markets in tourism.
method Developed and applied Bayesian Dirichlet autoregressive moving average (BDARMA) models to Airbnb booking data.
result BDARMA models outperform standard benchmarks in forecasting guest origin market shares.
Improved diffusion models for image synthesis with better training dynamics.
problem Uneven and ineffective training in diffusion models.
method Redesigned network layers to preserve activation, weight, and update magnitudes.
result Significantly better networks at equal computational complexity, improving FID to 1.81.
A function is exponentially concave if its exponential is concave. We consider exponentially concave functions on the unit simplex. In a previous paper we showed that gradient maps of exponentially concave functions provide solutions to a Monge-Kantorovich optimal transport problem and give a better gradient approximat…
Study compares local and global models for hierarchical forecasting accuracy.
problem Challenges in hierarchical time series forecasting, especially in accuracy and information utilisation.
method Developed and evaluated local and global forecasting models (GFMs) to exploit cross-series and cross-hierarchies information.
result Global Forecasting Models (GFMs) outperform local models in hierarchical forecasting accuracy and computational efficiency.
Gordian complex of knots was defined by Hirasawa and Uchida as the simplicial complex whose vertices are knot isotopy classes in S3. Later Horiuchi and Ohyama defined Gordian complex of virtual knots using v-move and forbidden moves. In this paper we discuss Gordian complex of knots by region crossing cha…
We show that any two geometric triangulations of a closed hyperbolic, spherical or Euclidean manifold are related by a sequence of Pachner moves and barycentric subdivisions of bounded length. This bound is in terms of the dimension of the manifold, the number of top dimensional simplexes and bound on the lengths of ed…
Positive weights improve kernel quadrature's accuracy.
problem Improving kernel quadrature weights to be positive and stable.
method Using convex geometry to approximate the kernel mean embedding with positive weights.
result Positive weights lead to improved kernel quadrature bounds with Monte-Carlo-beating rates.
Adaptive estimation of alpha-Stable distribution and Hurst exponent for nonstationary time series.
problem Nonstationary time series require adaptive models to avoid bias.
method Moving estimator with exponentially weakening weights of old values, optimized using EMA of absolute central moments.
result Continuous adaptive estimation of alpha-Stable distribution and Hurst exponent for market stability evaluation.