ATLAS adapts HMC step size and trajectory length for complex geometries.
problem Sampling complex geometries with constant step size HMC/NUTS.
method Adapts step size and trajectory length using local Hessian and no U-turn condition.
result ATLAS accurately samples complex geometries, outperforming NUTS.
Robust and fast method for large-scale stochastic optimization.
problem Large-scale stochastic optimization problems.
method Auxiliary variable construction coupled with adaptive inverse Hessian approximation.
result Encouraging performance on real-world problems with millions of observations and unknowns.
The CSA-ES is an Evolution Strategy with Cumulative Step size Adaptation, where the step size is adapted measuring the length of a so-called cumulative path. The cumulative path is a combination of the previous steps realized by the algorithm, where the importance of each step decreases with time. This article studies …
GIST adapts HMC by tuning parameters based on position and momentum.
problem Locally adaptive sampling in Hamiltonian Monte Carlo.
method GIST uses Gibbs sampling to adaptively tune HMC parameters.
result GIST improves sampling efficiency for high-dimensional models.
Adaptive TFTs improve cryptocurrency price prediction accuracy.
problem Precise short-term price prediction in volatile cryptocurrency markets.
method Dynamic subseries lengths and pattern-based categorization.
result Significantly outperforms baseline models in prediction accuracy and profitability.
Length spectral rigidity is the question of under what circumstances the geometry of a surface can be determined, up to isotopy, by knowing only the lengths of its closed geodesics. It is known that this can be done for negatively curved Riemannian surfaces, as well as for negatively-curved cone surfaces. Steps are tak…
Proposes a new algorithm for solving optimization problems with stochastic objectives and equality constraints.
problem Optimization problems with stochastic objectives and deterministic equality constraints.
method Trust-region stochastic sequential quadratic programming (TR-StoSQP) with adaptive relaxation techniques.
result Established a global almost sure convergence guarantee for TR-StoSQP.
Bayesian time series forecasting improves by dynamically adapting to recent information.
problem Lack of forgetting mechanism in signature kernel for time series forecasting.
method Introducing a novel forgetting mechanism for signature features using Random Fourier Decayed Signature Features (RFDSF) with Gaussian processes (GPs).
result Demonstrates superior performance compared to other GP-based alternatives and state-of-the-art probabilistic time series forecasting algorithms.
Study optimizes prediction intervals in conformal regression.
problem Optimizing the length of prediction intervals in conformal regression.
method Introduces EffOrt and Ad-EffOrt methodologies to minimize interval length.
result Demonstrates theoretical and empirical improvements over classical methods.
Two-Tailed Averaging improves generalization by optimizing the number of leading iterates to ignore.
problem Improving generalization in stochastic optimization with limited resources and hyperparameters.
method An anytime adaptive algorithm that balances the number of leading iterates to ignore for better generalization.
result Approximates the optimal tail at all optimization steps, improving generalization without hyperparameters.
Adaptive sampling improves graph diffusion models by maintaining uniform information speed.
problem Standard diffusion models overlook non-homogeneous dynamics on complex manifolds.
method Information-geometric framework using Fisher-Rao metric and Drift Variation Score (DVS).
result DVS solver ensures uniform rate of distributional change, improving structural fidelity and efficiency.
Proves C 1 C^1 C 1 regularity for abnormal minimizers in rank 2 sub-Riemannian structures.
problem Regularity of abnormal minimizers in sub-Riemannian structures.
method Proves C 1 C^1 C 1 regularity using length-minimizers in rank 2 sub-Riemannian structures. result All length-minimizers for rank 2 sub-Riemannian structures of step up to 4 are of class C 1 C^1 C 1 . Model learns brevity by exposing to easy problems, improving efficiency without explicit length penalties.
problem Excessive verbosity in step-by-step reasoning models trained with RLVR.
method Retaining and up-weighting moderately easy problems as implicit length regularizers.
result Model generates solutions that are, on average, nearly twice as short without explicit length penalties.
A new multi-step model improves model-based reinforcement learning efficiency.
problem Expensive environmental interaction in reinforcement learning.
method Proposes a multi-step model for predicting action sequences with variable length.
result Multi-step model outperforms one-step model in preliminary tests.
Introduces a new length functional for Ricci flow to detect steady solitons.
problem Detecting steady solitons in Ricci flow.
method Develops a modified length functional and shows it satisfies differential inequalities.
result The length functional generates a distance function that saturates on steady soliton manifolds.
Adaptive IHS improves sketching for large-scale data.
problem Efficiently modeling large-scale data with iterative Hessian sketch.
method Deterministic A-optimal subsampling for improved IHS.
result A-optimal IHS outperforms existing accelerated IHS methods.
Improved path-length regret bounds for adaptive and oblivious adversaries.
problem Adaptive and oblivious adversaries in multi-armed bandit and linear bandit problems.
method Developed two new algorithms based on optimistic mirror descent framework with novel techniques.
result Strictly improved path-length bounds for adaptive adversary and better results for oblivious adversary.
Adaptive TBPTT controls gradient bias in RNNs for faster convergence.
problem Choosing optimal truncation length in TBPTT for RNNs is difficult.
method Adaptive TBPTT converts lag selection to bias control, estimating optimal truncation length during training.
result Adaptive TBPTT improves convergence rate and computational efficiency in RNNs.
The subject of this paper is the relationship among the marked length spectrum, the length spectrum, the Laplace spectrum on functions, and the Laplace spectrum on forms on Riemannian nilmanifolds. In particular, we show that for a large class of three-step nilmanifolds, if a pair of nilmanifolds in this class has the …
The subject of this paper is the relationship among the marked length spectrum, the length spectrum, the Laplace spectrum on functions, and the Laplace spectrum on forms on Riemannian nilmanifolds. In particular, we show that for a large class of three-step nilmanifolds, if a pair of nilmanifolds in this class has the …
Proposes RN for unsupervised attention in neural networks.
problem Limited, imbalanced, and non-stationary input distributions in various tasks.
method Inspired by neuronal adaptation, RN uses MDL principle and universal code length for incremental layer-wise computation.
result Outperforms existing normalization methods across diverse tasks.
Study gluing of Lorentzian length spaces and their causal ladder properties.
problem Compatibility of Lorentzian amalgamation with length space properties.
method Conditions for gluing Lorentzian length spaces and criteria for causal ladder preservation.
result Gluing of Lorentzian length spaces yields again a Lorentzian length space under certain conditions.
MB-DQN uses different backup lengths for improved reinforcement learning.
problem Improving reinforcement learning with multi-step returns.
method Integrates multi-step returns into bootstrapped DQN with different backup lengths.
result MB-DQN maintains advantages of different backup lengths.
AdaSVRG combines adaptive gradient with SVRG for robust optimization.
problem Variance reduction in finite-sum minimization with unknown constants.
method AdaSVRG uses AdaGrad in SVRG's inner loop to make it robust.
result AdaSVRG achieves optimal gradient evaluations with no need for problem-dependent constants.
PROFET builds DBNs from ODEs, handling uncertainty in data and models.
problem Handling uncertainty in both data and models for DBN construction.
method Automatic DBN construction from ODEs, adaptive-time particle filtering.
result PROFET automates DBN construction and inference from ODE models.
Two adaptive algorithms improve tracking regret in dynamic expert advice problems.
problem Prediction with expert advice in dynamic environments.
method Developed two adaptive and efficient algorithms using online mirror descent framework.
result Achieved data-dependent tracking regret bounds for both algorithms.
An adaptive time-stepping controller improves stability and accuracy of ResNets.
problem Improving stability and performance of ResNets using adaptive time stepping.
method Developed an adaptive time-stepping controller based on Runge-Kutta-Fehlberg method.
result Demonstrated improved stability and accuracy of ResNets without additional overhead.
TIDBD adapts TD step-sizes for better performance.
problem Finding optimal step-sizes for TD learning.
method Generalizes IDBD to TD learning, adapting step-sizes per feature.
result TIDBD outperforms other TD methods in various tasks.
A new BO method adapts hyperparameters online and uses a novel kernel for global and local optimization.
problem Expensive black-box optimization problems.
method Online length-scale adaption, mixed-global-local kernel, and adaptive hyperparameters.
result The proposed method outperforms state-of-the-art BO methods on global optimization benchmarks.
Adaptive step-size improves optimization in complex geometries.
problem Optimizing functions with non-Euclidean geometries.
method Adaptive step-size strategy for optimization algorithms.
result Guaranteed convergence for Adaptive Conditional Gradient Descent.
New adaptive step-size method for convex optimization without tuning.
problem Optimizing convex functions efficiently with stochastic gradients.
method Adapted Adaptive Gradient Descent Without Descent to stochastic setting.
result Stochastic gradient descent converges under various assumptions.
Paper adapts ACI for online multi-step time-series forecasting with coverage guarantees.
problem Achieving reliable error bounds in online multi-step time-series forecasting.
method Adaptive conformal inference (ACI) adapted for multi-step forecasting with dynamic significance levels.
result Proposes a multi-step ACI algorithm with finite-sample coverage guarantees for non-exchangeable data.
New SVRG and SARAH schemes reduce tuning effort for variance reduction.
problem Optimal performance of SVRG and SARAH requires tuning of parameters.
method Introduces Barzilai-Borwein step sizes, averaging, and adaptive inner loop length.
result Improves SVRG, SARAH, and BB variants' convergence rates and performance.
MELO predicts electricity loads by adapting to shifts without external indicators.
problem Adapting to non-stationary prediction challenges in online settings.
method MELO combines multiple forgetting factors and aggregation rules to adaptively predict.
result MELO reduces RMSE by 34.7% compared to base predictors and external covariates.
Random walk on hyperbolic surfaces yields uniform distribution of hyperbolic elements.
problem Distribution of lengths in random walks on Fuchsian groups.
method Proof using Gromov's theorem on translation lengths of Gromov-hyperbolic groups.
result Geometric lengths follow laws of large numbers, central limit theorem, etc.
Transformer pretraining yields strong EB performance without explicit adaptation.
problem Empirical Bayes problems with unknown test distributions.
method Indirect analysis of pretrained transformer's performance under universal priors.
result Near-optimal regret bound of O ~ ( 1 n ) \widetilde{O}(\frac{1}{n}) O ( n 1 ) for arbitrary test distributions. New online conformal prediction methods minimize strongly adaptive regret and achieve near-optimal coverage.
problem Uncertainty quantification in online settings with changing data distributions.
method Developed new online conformal prediction methods that minimize strongly adaptive regret.
result Achieve near-optimal strongly adaptive regret and approximately valid coverage.
The paper introduces a new method to infer phylogenetic trees without bifurcations.
problem Inferring phylogenetic trees with zero-length branches and polytomies.
method Adaptive LASSO-type regularization estimators for phylogenetics.
result Regularization is a practical approach for phylogenetics, revealing zero-length branches.
New convergence analysis for ADAM algorithm in non-convex optimization with adaptive step size.
problem Convergence issues in ADAM algorithm for non-convex optimization.
method Study of ADAM algorithm under bounded adaptive step size assumption, providing safe step sizes.
result Novel first order convergence rate result in deterministic and stochastic contexts.
Adaptive step sizes improve optimization for convex and nonconvex problems.
problem Optimizing functions that are not strongly convex.
method Bridge nonconvex and strongly convex problems via regularization, then apply Barzilai-Borwein step sizes with SARAH.
result Regularized SARAH methods achieve better complexity in nonconvex problems.
Adaptive step-size method improves compressed SGD performance in machine learning.
problem Communication bottleneck in distributed and decentralized optimization.
method Developed an adaptive step-size method for compressed SGD.
result Order-optimal convergence rates for various objective functions.
New approach reduces malware detection memory requirements and speeds up training.
problem Efficiently classifying long sequences of malware detection data.
method Developed a new temporal max pooling method and global channel gating design.
result 116x more memory efficient and 25.8x faster training on original dataset.
Paper proposes an active multi-step TD algorithm for reinforcement learning.
problem Challenging decision making and control tasks in reinforcement learning.
method Active stepsize learning and adaptive multi-step TD algorithm with context-aware mechanism.
result Competitive results compared to other reinforcement learning baselines on discrete and continuous space tasks.
BCI provides calibrated prediction intervals for time series forecasts.
problem Calibration of prediction intervals for time series forecasts.
method BCI wraps around any time series forecasting models and optimizes interval lengths using dynamic programming.
result BCI achieves long-term coverage under arbitrary distribution shifts and temporal dependence.
The paper analyzes and validates two step size schedules for SGD: exponential and cosine, proving their adaptivity and performance.
problem The variability of SGD performance due to step size choice.
method Analysis and empirical evaluation of exponential and cosine step sizes.
result Exponential and cosine step sizes are adaptive to noise and achieve optimal performance without tuning hyperparameters.
RATQ is a new quantizer for optimizing noisy gradients in machine learning.
problem Optimizing noisy gradients in stochastic optimization.
method RATQ uses Hadamard transform and adaptive uniform quantization, and achieves near-optimal performance.
result RATQ nearly achieves information theoretic lower bounds for optimization accuracy.
A new method improves convergence in large-scale stochastic optimisation.
problem Improving convergence in large-scale stochastic optimisation problems.
method A direct least-squares approach with a Cholesky factor and adaptive line search.
result Improved convergence compared to existing methods on real-world problems.
Step-DAD improves BED by periodically updating a design policy during experiments.
problem Improving flexibility and robustness in Bayesian experimental design.
method Semi-amortized, policy-based approach that updates a design policy during data collection.
result Consistently superior decision-making and robustness compared to current BED methods.