Improved stability for large-scale Bayesian sampling.
problem Reducing instability in Langevin dynamics for large datasets.
method Introducing a modified CCAdL thermostat with a scaling and squaring method and a truncated Taylor series approximation.
result Significantly improved numerical stability and accuracy over existing methods.
The study shows conditions for thermostats to have no conjugate points.
problem Conditions for thermostats to have no conjugate points.
method Analyzes smooth functions and properties of thermostats.
result Proves conditions for thermostats to be projectively Anosov and have no conjugate points.
Monte Carlo sampling for Bayesian posterior inference is a common approach used in machine learning. The Markov Chain Monte Carlo procedures that are used are often discrete-time analogues of associated stochastic differential equations (SDEs). These SDEs are guaranteed to leave invariant the required posterior distrib…
The paper studies ray transforms on surfaces with negative curvature, proving injectivity and determining connections and Higgs fields.
problem Injectivity of ray transforms on surfaces with negative curvature and determination of connections and Higgs fields.
method Analysis of Gaussian thermostats on compact Riemannian surfaces with negative curvature, proving injectivity and determining connections and Higgs fields.
result Injectivity of the thermostat ray transform and determination of connections and Higgs fields.
In this paper we consider the Gaussian thermostat ray transform on both closed Riemannian surfaces and compact Riemannian surfaces with boundary. We establish certain results on the injectivity of the thermostat ray transform and the surjectivity of its adjoint.
Smooth orbit equivalence proves metric equivalence for geodesic flows.
problem Proving metric equivalence for geodesic flows under orbit equivalence.
method Proving metric equivalence for geodesic flows under orbit equivalence.
result Smooth orbit equivalence implies conformal equivalence of metrics.
We show a Hopf type rigidity for thermostats without conjugate points on a 2-torus
We propose a new sampling method, the thermostat-assisted continuously-tempered Hamiltonian Monte Carlo, for Bayesian learning on large datasets and multimodal distributions. It simulates the Nosé-Hoover dynamics of a continuously-tempered Hamiltonian system built on the distribution of interest. A significant advantag…
Guillarmou extends X-ray transform to magnetic and thermostat flows.
problem Stability of magnetic X-ray transforms.
method Generalizes normal operator to thermostat and magnetic flows, proving ellipticity.
result Elliptic pseudodifferential operators of order -1 for generalized normal operators.
In this paper, we will recover Hamilton's Harnack inequality for the Ricci flow from the view point of Hyperbolic thermostat.
Researchers study injectivity of magnetic and thermostatic nonabelian ray transforms on compact surfaces.
problem Injectivity of magnetic and thermostatic nonabelian ray transforms on compact surfaces.
method Loop group factorization method for nontrapping λ-geodesic flows and the general linear group of invertible complex matrices. result General injectivity question of the nonabelian ray transform for simple magnetic flows is settled.
Alternative construction of quasi-Fuchsian flows using vortex equations.
problem Constructing quasi-Fuchsian flows.
method Using coupled vortex equations to construct quasi-Fuchsian flows as thermostats.
result Formulas for the marked length spectrum of quasi-Fuchsian flows.
RL applied to TCLs for power consumption control.
problem Optimizing power consumption using TCLs with RL.
method Modelica-based reinforcement learning (Q-learning) for stochastic TCLs.
result Q-learning parameters affect controller performance.
Learning in deep models using Bayesian methods has generated significant attention recently. This is largely because of the feasibility of modern Bayesian methods to yield scalable learning and inference, while maintaining a measure of uncertainty in the model parameters. Stochastic gradient MCMC algorithms (SG-MCMC) a…
We introduce a new family of thermostat flows on the unit tangent bundle of an oriented Riemannian 2-manifold. Suitably reparametrised, these flows include the geodesic flow of metrics of negative Gauss curvature and the geodesic flow induced by the Hilbert metric on the quotient surface of divisible convex sets. We …
We consider a general family of curves Γ on a compact oriented Finsler surface (M,F) with boundary ∂M. Let φ∈C∞(M) and ω a smooth 1-form on M. We show that ∫γ(t){φ(γ(t))+ωγ(t)(γ˙(t))}dt=0 holds for every γ∈Γ whose endpoints belong to ∂M, $γ(a)…
Recent studies have shown that the aggregated dynamic flexibility of an ensemble of thermostatic loads can be modeled in the form of a virtual battery. The existing methods for computing the virtual battery parameters require the knowledge of the first-principle models and parameter values of the loads in the ensemble.…
QHMC improves HMC for sampling from complex distributions.
problem Inefficiency of HMC in sampling from spiky and multimodal distributions.
method Proposes QHMC, a quantum-inspired version of HMC with a random mass matrix.
result QHMC and QSGNHT achieve more stable and accurate sampling results.
We use online convex optimization (OCO) for setpoint tracking with uncertain, flexible loads. We consider full feedback from the loads, bandit feedback, and two intermediate types of feedback: partial bandit where a subset of the loads are individually observed and the rest are observed in aggregate, and Bernoulli feed…
A new hierarchy quantifies agency in systems based on information processing.
problem Lack of a measurable, universal definition for agency in intelligent systems.
method Developed a bottom-up framework based on information processing hierarchy.
result Identified three orders of information processing (I, II, III) as necessary for agency.
Study proves uniqueness for ray transform on surfaces with obstacles.
problem Uniqueness of functions and 1-forms on surfaces with reflecting obstacles.
method Broken ray transform on twisted geodesics with nonpositive curvature and reflecting boundary.
result Proves uniqueness result for sums of functions and 1-forms.
Study resonant forms for dissipative Anosov flows on 3-manifolds.
problem Determine resonant forms and their cohomology classes for dissipative Anosov flows.
method General theory including horocyclic invariance and local geometry analysis.
result Explicit computation of resonant forms and helicity for quasi-Fuchsian flows.
Bayesian max-margin models have shown superiority in various practical applications, such as text categorization, collaborative prediction, social network link prediction and crowdsourcing, and they conjoin the flexibility of Bayesian modeling and predictive strengths of max-margin learning. However, Monte Carlo sampli…
Recent advances in Bayesian learning with large-scale data have witnessed emergence of stochastic gradient MCMC algorithms (SG-MCMC), such as stochastic gradient Langevin dynamics (SGLD), stochastic gradient Hamiltonian MCMC (SGHMC), and the stochastic gradient thermostat. While finite-time convergence properties of th…
Proposes a method to handle missing inputs in Bayesian optimization.
problem Missing values in historical data and function evaluations.
method Impute missing values using probability distributions and develop a new acquisition function.
result Improves performance of Bayesian optimization by handling missing inputs effectively.
Paper tackles adaptive submodularity in sequential decision making.
problem Maximizing adaptive submodular functions with limited adaptive rounds.
method Proposes an efficient semi-adaptive policy with logarithmic adaptive rounds.
result Achieves an almost tight 1−1/e−ε approximation guarantee. Stochastic methods with coordinate-wise adaptive stepsize (such as RMSprop and Adam) have been widely used in training deep neural networks. Despite their fast convergence, they can generalize worse than stochastic gradient descent. In this paper, by revisiting the design of Adagrad, we propose to split the network par…
Sequence-to-sequence (seq2seq) based ASR systems have shown state-of-the-art performances while having clear advantages in terms of simplicity. However, comparisons are mostly done on speaker independent (SI) ASR systems, though speaker adapted conventional systems are commonly used in practice for improving robustness…
AdaPTS adapts univariate FMs for multivariate time series forecasting.
problem Challenges in managing feature dependencies and uncertainty quantification in multivariate time series forecasting.
method Adapters that transform multivariate inputs into a latent space and apply univariate FMs independently to each dimension.
result AdaPTS enhances forecasting accuracy and uncertainty quantification compared to baseline methods.
We propose a new concept named adaptive submodularity ratio to study the greedy policy for sequential decision making. While the greedy policy is known to perform well for a wide variety of adaptive stochastic optimization problems in practice, its theoretical properties have been analyzed only for a limited class of p…
New algorithm reduces interventional strategy complexity for causal graph discovery.
problem Designing efficient interventional strategies for causal graph discovery.
method Developed an r-adaptive algorithm for causal graph discovery that minimizes the number of interventions.
result Achieved an approximation of O(min{r, log n} * n^{1/min{r, log n}}) for the verification number.
Automation of machine learning model development is increasingly becoming an established research area. While automated model selection and automated data pre-processing have been studied in depth, there is, however, a gap concerning automated model adaptation strategies when multiple strategies are available. Manually…
Paper presents a Transformer model for automatic domain adaptation.
problem Challenges in selecting or designing domain adaptation algorithms.
method Transformer model approximates and selects domain adaptation algorithms.
result Transformers can approximate and automatically select domain adaptation algorithms.
cKAM improves adaptive sampling by incorporating a cyclical stepsize scheme.
problem Adaptive Metropolis algorithms can get stuck in local modes.
method cKAM uses a cyclical stepsize scheme to encourage exploration and escape from local modes.
result cKAM successfully escapes local modes and converges to the true posterior distribution.
Adaptive variational Bayes framework improves inference adaptively.
problem Lack of general and computationally tractable variational Bayes method for adaptive inference.
method Proposes a novel adaptive variational Bayes framework combining variational posteriors over individual models.
result Adaptive variational Bayes achieves optimal contraction rates adaptively under general conditions.
New approach shows AI can adapt like toddlers by correcting old knowledge.
problem Lack of adaptability in AI models compared to humans and animals.
method Casting adaptation as posterior correction and using Bayesian Learning Rule.
result AI can learn to adapt quickly by using posterior correction.
This paper improves neural network generalization by dynamically learning kernel parameters.
problem Improving neural network generalization and adaptability.
method Diagonal adaptive kernel model that learns kernel eigenvalues and output coefficients during training.
result The diagonal adaptive kernel model significantly improves generalization over fixed-kernel methods.
Study on distributed nonparametric function estimation with optimal rate and cost of adaptation.
problem Optimal rate of convergence and cost of adaptation in distributed nonparametric function estimation.
method Distributed minimax estimation and adaptive estimation under communication constraints for Gaussian sequence model and white noise model.
result Established minimax rate of convergence and exact communication cost for adaptation.
This paper explains why Adam generalizes worse than SGD by analyzing its components.
problem Understanding why Adam generalizes worse than Stochastic Gradient Descent (SGD).
method Diffusion theoretical framework to disentangle the effects of Adaptive Learning Rate and Momentum.
result Adaptive Learning Rate helps escape saddle points but not select flat minima, while Momentum provides a drift effect to help pass through saddle points.
New adaptive importance samplers improve stability and accuracy.
problem Improving the stability and accuracy of importance sampling estimators.
method Introducing AdaOAIS, a new adaptive importance sampler using adaptive optimisers to address the instability of OAIS.
result AdaOAIS leads to stable importance sampling estimators in practice.
Adaptive networks improve model robustness through conditional normalization.
problem Limited robustness of adversarial-trained networks due to network capacity and training samples.
method Proposes a conditional normalization module to adapt networks during adversarial training.
result Adaptive networks outperform both clean validation accuracy and robustness compared to non-adaptive counterparts.
In domain adaptation, classifiers with information from a source domain adapt to generalize to a target domain. However, an adaptive classifier can perform worse than a non-adaptive classifier due to invalid assumptions, increased sensitivity to estimation errors or model misspecification. Our goal is to develop a doma…
FLAP adapts policies quickly to new tasks using shared linear representations.
problem Adapting policies to new tasks efficiently and effectively.
method FLAP uses a shared linear representation and a separate adapter network for quick adaptation.
result FLAP achieves up to 8X faster adaptation and significantly better performance on out-of-distribution tasks.
AvaGrad optimizes vision tasks by decoupling learning rate and adaptability.
problem Improving optimization methods for vision tasks.
method Derives AvaGrad, a new optimizer that decouples learning rate and adaptability.
result AvaGrad outperforms SGD on vision tasks when adaptability is properly tuned.
Learn to automatically plug domain-specific modules into a common network.
problem Learning inflexibility and computational intensiveness in multi-domain learning.
method Neural Architecture Search (NAS) for data-driven adapter plugging and structure design.
result NAS-driven MDL model achieves comparable performance to existing approaches.
FLoE adapts LLMs by selectively deploying LoRA adapters based on layer importance and task requirements.
problem Uniform LoRA deployment across all layers leads to inefficient and redundant parameter allocation.
method FLoE uses Fisher information to dynamically identify task-critical layers and optimizes LoRA ranks.
result FLoE achieves significant efficiency-accuracy trade-offs, especially in resource-constrained environments.
The paper improves generalization bounds for domain adaptation.
problem Improving generalization bounds for domain adaptation under practical conditions.
method Derives generalization bounds for domain adaptation based on finitely many moments and smoothness conditions.
result Obtains generalization bounds for domain adaptation.
New protocols show 1-bit mean estimation can be order-optimal without interaction.
problem Can 1-bit mean estimation be optimal without interaction?
method Adaptive and non-adaptive threshold and interval queries, with one adaptive transition.
result Arbitrary non-adaptive quantizers can match the adaptive rate, suggesting interaction is not necessary.