Generative model MSM improves on handwritten multidialect data.
problem Inefficient representation of multidialect handwritten data.
method Two-point error update for hierarchical mode synthesis.
result MSM achieves lower error values than RBM on independent and mixed data.
The study examines conditions for achieving a simple lower bound in estimating mean from samples.
problem Achieving a simple lower bound for estimating the mean of a distribution.
method Analyzes conditions for nearly attaining Le Cam's two-point testing lower bound for mean estimation.
result An algorithm nearly attains the two-point testing rate for mixtures of symmetric, log-concave distributions with a common mean.
Quandles can be regarded as generalizations of symmetric spaces. Among symmetric spaces, two-point homogeneous Riemannian manifolds would be the most fundamental ones. In this paper, we define two-point homogeneous quandles analogously, and classify those with prime cardinality.
SGD reduces test error by decorrelating updates.
problem Improving generalization error in machine learning models.
method Derive a formula for generalization gap change due to SGD updates, compare to GD, and show decorrelation effect.
result SGD implicitly regularizes generalization error by decorrelating updates.
Optimal algorithm for bandit and zero-order convex optimization with two-point feedback.
problem Optimal algorithm for convex optimization with two-point feedback.
method Simple algorithm based on a modified gradient estimator.
result Optimal for convex Lipschitz functions, improving on previous results for smooth functions.
The paper bounds generalization error for iterative learning with bounded updates.
problem Generalization error of iterative learning algorithms with bounded updates for non-convex loss functions.
method Information-theoretic techniques, reformulating mutual information as update uncertainty, variance decomposition.
result Improved generalization error bounds for iterative learning algorithms with bounded updates.
Let G be the identity component of the isometry group for an arbitrary curved two-point homogeneous space M. We consider algebras of G-invariant differential operators on bundles of unit spheres over M. The generators of this algebra and the corresponding relations for them are found. The connection of these ge…
SUMER updates deployed models with new data, improving performance.
problem Updating deployed machine learning models with new data is difficult.
method SUMER uses semi-supervised learning and noise remediation to iteratively retrain models.
result SUMER improves model performance, especially with limited initial training data.
In this paper we show that on a complete Riemannian manifold of negative curvature and dimension n>1 every two points which realize a local maximum for the distance function are connected by at least 2n+1 geometrically distinct geodesic segments (i.e. length minimizing). Using a similar method, we obtain that in th…
Improved HGF networks avoid negative precision errors in volatility updates.
problem Negative posterior precision errors in volatility-coupled nodes of HGF networks.
method Introduced a modified quadratic approximation to variational energy.
result Robust update equations across parameter space that track posterior faithfully.
Two-root Riemannian manifolds have no odd-dimensional examples.
problem Characterizing Riemannian manifolds with specific eigenvalues of the Jacobi operator.
method Investigation of k-root manifolds, focusing on one-root and two-root cases. result There are no two-root Riemannian manifolds of odd dimension.
Using a two-point correlation technique, we study emergence of market efficiency in the emergent Russian futures market by focusing on lagged correlations. The correlation strength of leader-follower effects in the lagged inter-market correlations on the hourly time frame is seen to be significant initially (2009-2011)…
Estimates for solutions on manifolds under Ricci flow.
problem Gradient estimates for solutions on manifolds.
method Two-point function estimates for quasilinear parabolic equations under Ricci flow.
result Gradient estimates for solutions at any two points related to the distance between points.
Proves convexity of minimizers in energy functions with convex potentials.
problem Connectedness and convexity of minimizers in energy functions involving surface tensions and convex potentials.
method Introduces a 'two-point function' to measure lack of convexity and prove negative second variation of the energy.
result Positively answers an old question of Almgren about connectedness and convexity of minimizers.
The article completes the research of two-point G2 Hermite interpolation problem with spirals by inversion of conics. A simple algorithm is proposed to construct a family of 4th degree rational spirals, matching given G2 Hermite data. A possibility to reduce the degree to cubic is discussed.
Enhances linear regression with Kalman filter for loss minimization.
problem Minimizing loss in linear regression models.
method Integrates Kalman filter and SGD for optimal weight updates.
result Develops optimal linear regression equation with minimum area under curve.
New bounds for noisy iterative algorithms reduce overfitting risk.
problem Bounding generalization error for noisy, iterative algorithms.
method Derive bounds on generalization error using mutual information and sub-Gaussian loss function.
result Generalization error bounds for a broad class of iterative algorithms.
Smooth solutions found for hydrodynamic equations.
problem Geodesic equations on diffeomorphism groups.
method Proving smoothness of solutions and constructing exponential maps.
result Smooth solutions match boundary conditions.
The study examines the growth of conjugacy classes in negatively curved manifolds.
problem Growth of conjugacy classes in manifolds with variable negative curvature.
method Analyzes the number of conjugacy class orbits in a ball of radius T centered at a point in the universal cover of a manifold.
result Exponentially small error terms for the count of conjugacy class orbits in 2D or high-dimensional manifolds with specific curvature bounds.
The paper addresses Dyna-style RL's value hallucination issue by proposing a new algorithm.
problem Value hallucination in Dyna-style RL due to bootstrapping simulated states.
method Introduces a new Dyna algorithm using predecessor models with multi-step updates.
result Evidence supports the Hallucinated Value Hypothesis (HVH), suggesting predecessor models with multi-step updates are promising.
New method reduces overestimation in actor-critic reinforcement learning.
problem Function approximation errors in actor-critic methods lead to suboptimal policies.
method Proposes novel mechanisms to minimize overestimation, including using the minimum value between critics and delaying policy updates.
result Outperforms state-of-the-art methods on OpenAI gym tasks.
Novel method for nonlinear data assimilation using Langevin sampling.
problem Nonlinear data assimilation challenges in Bayesian filtering.
method Score-based sequential Langevin sampling (SSLS) with dynamic models and annealing.
result Asymptotic stability and error bounds for local posterior sampling.
We consider the two body problem with central interaction on two point homogeneous spaces from point of view of the invariant differential operators theory. The representation of the two particle Hamiltonian in terms of the radial differential operator and invariant operators on the symmetry group is found. The connect…
Training-free source selection for LLM families with shared vocabularies
problem Source selection for LLM families with shared vocabularies
method Fisher alignment at vocabulary scale
result Fisher alignment is a cosine between kernel mean embeddings in the joint activation-error space
Study on distributed coordinate descent with quantized updates for finite precision communication.
problem Finite precision communication limits the accuracy of updates in distributed coordinate descent.
method Introduced a randomized distributed coordinate descent algorithm with quantized updates, derived convergence conditions, and validated with experiments.
result Algorithm with quantized updates converges under certain conditions on the quantization error.
AdaComm optimizes SGD by dynamically adjusting communication frequency for faster convergence.
problem Achieving optimal error-runtime trade-off in distributed SGD.
method Adaptive communication strategy that starts with infrequent averaging to save delay and improve speed, then increases frequency.
result AdaComm reduces training time by 3x while maintaining the same final loss.
AdaQuantFL reduces communication in federated learning by adaptively quantizing model updates.
problem Efficient communication of model updates in federated learning with high-dimensional models and limited bandwidth.
method AdaQuantFL uses adaptive quantization to reduce the number of bits for model updates while maintaining low error floor.
result AdaQuantFL converges in fewer communicated bits compared to fixed quantization levels, with minimal impact on accuracy.
Proposes a new method for nonlinear Bayesian updates using ensemble kernel regression.
problem Nonlinear and non-Gaussian Bayesian updates for complex systems.
method Combines Kalman filtering for observed components and kernel density estimation for unobserved components, with subsampling and clustering.
result Reduces estimation errors in highly nonlinear scenarios compared to standard linear updates.
Study MAML's generalization in varying tasks, proving bounds on error.
problem Bounding MAML's generalization error across tasks.
method Characterizes MAML's generalization error from two perspectives: recurring and unseen tasks.
result MAML's generalization error depends on the number of tasks and samples per task.
Study on PG learning for LQ MFC problems with common noise, proving convergence and sample complexity.
problem Optimal policy learning in LQ MFC problems with common noise and entropy regularization.
method Comprehensive error analysis of PG algorithms in both model-based and model-free settings.
result Global linear convergence and sample complexity of PG algorithms in model-free setting.
Selective state-adaptive regularization improves offline RL performance.
problem Extrapolation errors and value overestimation in static dataset RL.
method State-adaptive regularization coefficients trust Bellman-driven results selectively.
result Significant improvement in performance on D4RL benchmark.
Model forgets examples; this research predicts which ones to replay.
problem Language models forget examples during updates, leading to errors.
method Train forecasting models to predict which examples will be forgotten.
result Forecasting models can reduce forgetting of upstream pretraining examples.
Mixed labyrinth fractals can have finite or infinite arc lengths.
problem Characterizing the length of arcs in mixed labyrinth fractals.
method Analyzing sequences of labyrinth patterns to determine arc lengths.
result Arc lengths can be finite or infinite depending on pattern choice.
Unified derivation of high-dimensional linear models using stochastic gradient descent.
problem Performance analysis of high-dimensional linear models trained with stochastic gradient descent.
method Derivation of a deterministic equivalence for the two-point function of a random matrix resolvent.
result Unified understanding of model performance including previously known and novel results.
Let g be a metric on S3 with positive Yamabe constant. When blowing up g at two points, a scalar flat manifold with two asymptotically flat ends is produced and this manifold will have compact minimal surfaces. We introduce the $\Th$-invariant for g which is an isoperimetric constant for the cylindrical domain…
A new one-point feedback scheme improves ZO algorithms for black-box optimization.
problem Optimizing black-box functions without gradient information.
method Proposes a one-point feedback scheme to estimate gradients using residuals.
result Matches query complexity of two-point schemes for deterministic Lipschitz functions.
This paper provides a block coordinate descent algorithm to solve unconstrained optimization problems. In our algorithm, computation of function values or gradients is not required. Instead, pairwise comparison of function values is used. Our algorithm consists of two steps; one is the direction estimate step and the o…
Study shows properties of noncompact hypersurfaces in hyperbolic space.
problem Characterize noncompact hypersurfaces in hyperbolic space with nonnegative Ricci curvature.
method Utilized properties of n-subharmonic functions to analyze asymptotic boundaries.
result Hypersurfaces with nonnegative Ricci curvature in hyperbolic space have at most two points in their asymptotic boundary.
Paper introduces statistical learning for point processes.
problem Statistical learning for point processes in general spaces.
method Combines bivariate innovations and point process cross-validation.
result Statistical learning approach outperforms state of the art.
We investigate finite-time decoupled convergence in nonlinear two-time-scale stochastic approximation.
problem Achieving decoupled convergence in nonlinear two-time-scale stochastic approximation.
method Nested local linearity assumption, suitable step size selection, convergence analysis of matrix cross term, fourth-order moment convergence rates.
result Finite-time decoupled convergence rates can be achieved in nonlinear two-time-scale stochastic approximation with proper step size selection.
Bio-inspired neural networks use predictive coding for efficient weight updates.
problem Training artificial neural networks efficiently and biologically plausibly.
method Predictive Coding (PC) updates weights locally using only local information.
result PC provides theoretical advantages like automatic gradient scaling.
Efficiently merges multiple points to speed up BSGD SVM training.
problem Costly merging of points in BSGD SVM training.
method Merges more than two points at once to reduce training time.
result Significant speed-ups achieved without loss of accuracy.
A new method matches point sets of low-rank networks via their Laplace transforms.
problem Matching nodes in unseeded, low-rank networks without known correspondences.
method Transform-based unsupervised point registration via minimizing discrepancy between Laplace transforms.
result First consistency guarantee and explicit error rate for general low-rank models.
VCoTTA uses variational Bayesian methods to adapt models under continuous domain shifts.
problem Error accumulation in continual test-time adaptation.
method VCoTTA employs variational Bayesian techniques to update a Bayesian Neural Network (BNN) during testing, combining priors from source and teacher models.
result VCoTTA effectively mitigates error accumulation in CTTA, as shown by experimental results on three datasets.
We present a new online boosting algorithm for adapting the weights of a boosted classifier, which yields a closer approximation to Freund and Schapire's AdaBoost algorithm than previous online boosting algorithms. We also contribute a new way of deriving the online algorithm that ties together previous online boosting…
DEAM optimizes momentum weights dynamically to improve deep learning model training.
problem Errors in momentum weights propagate errors in optimization algorithms like ADAM.
method DEAM computes adaptive momentum weights based on discriminative angles, reducing hyperparameters and introducing a backtrack term.
result DEAM achieves faster convergence rates in both convex and non-convex deep learning model training.
Designing deterministic denominators for SGLD stabilizes large drifts.
problem Stabilizing large drifts in SGLD
method Using state-dependent envelopes and empirical quantiles for activation thresholds
result Proxy-quantile denominators are close to oracle-score behavior and improve deterministic taming choices
New algorithm accelerates single-pass SGD for generalized linear prediction.
problem Improving single-pass non-quadratic stochastic optimization.
method Data-dependent proximal method incorporating dual-momentum acceleration.
result Momentum acceleration resolves open problem in streaming setting.