FP-UCB algorithm achieves bounded regret for finitely parameterized multi-armed bandits.
problem Finitely parameterized multi-armed bandits with unknown but known parameter set.
method FP-UCB algorithm using structural information about the parameter set.
result FP-UCB achieves bounded regret under structural condition, logarithmic otherwise.
The paper studies multi-curve interest rate models and their consistency and finite-dimensional realizations.
problem Consistency and existence of finite-dimensional realizations for multi-curve interest rate models.
method Geometric approach, characterizing consistency and existence of finite-dimensional realizations for multi-curve models.
result Characterization of consistency and existence of finite-dimensional realizations for multi-curve models.
Learning optimal resource allocation policies in wireless systems can be effectively achieved by formulating finite dimensional constrained programs which depend on system configuration, as well as the adopted learning parameterization. The interest here is in cases where system models are unavailable, prompting method…
Improved standard parameterization yields well-defined neural tangent kernel.
problem Extrapolation of standard parameterization to infinite width is problematic.
method Proposed an improved extrapolation of the standard parameterization.
result Improved standard parameterization yields similar accuracy to NTK parameterization but with better correspondence to finite width networks.
Study calculates the elastic energy of curves on a sphere.
problem Elastic energy of curves on a sphere.
method Introduced p-curvature functional for rectifiable curves in the sphere and proved its finiteness. result The p-curvature functional agrees with the integral of geodesic curvature raised to the power p for curves in W2,p. This paper tackles constrained statistical learning problems by proposing a new approach.
problem Statistical learning problems with constraints are challenging and scarce.
method Directly tackling the constrained problem using finite dimensional parameterizations, sample averages, and duality theory.
result We bound the empirical duality gap, showing the effectiveness of the constrained formulation.
Kernel methods are powerful tools to capture nonlinear patterns behind data. They implicitly learn high (even infinite) dimensional nonlinear features in the Reproducing Kernel Hilbert Space (RKHS) while making the computation tractable by leveraging the kernel trick. Classic kernel methods learn a single layer of nonl…
The current paper discusses some new results about conformal polynomic surface parameterizations. A new theorem is proved: Given a conformal polynomic surface parameterization of any degree it must be harmonic on each component. As a first geometrical application, every surface that admits a conformal polynomic paramet…
Vogel's construction links knot invariants to Lie algebras, revealing new insights.
problem Can all finite type knot invariants be derived from Lie algebras?
method Parameterized expansion coefficients with three parameters and constructed a polynomial to vanish for all simple Lie algebras.
result Vogel's construction implies an alternative axiomatization of simple Lie algebras.
The paper proposes methods for volumetric parameterization of 3D solid manifolds.
problem Complex structure of solid manifolds makes conventional approaches ineffective.
method Incorporates models to preserve geometric structure, achieve density equalization, and balance distortions.
result Various 3D manifold parameterizations with different properties can be achieved.
Local PCA detects intrinsic parameterization of complex thermo-chemical state-spaces.
problem Detecting intrinsic parameterization of complex thermo-chemical state-spaces.
method Local PCA applied to local clusters of data.
result Local PCA finds meaningful parameterization linked to local stoichiometry, reaction progress, and soot formation processes.
Agents collaborate to reduce regret in a multi-agent linear bandit problem with side information.
problem Reducing regret in a multi-agent stochastic linear bandit with side information.
method A decentralized algorithm where agents communicate subspace indices and each plays a projected LinUCB on the corresponding low-dimensional subspace.
result Per-agent finite-time regret is much smaller when agents communicate compared to non-communicating case.
Minimal 7D hypersurfaces degenerate under stability or bounded index constraints.
problem Degeneration of minimal hypersurfaces under stability or bounded index constraints.
method Analysis of sequences of minimal hypersurfaces, parameterization with controlled maps, and topological finiteness results.
result Minimal hypersurfaces can degenerate to singular ones with controlled geometry, topology, and singular set.
The paper explains how ReLU nets converge globally in high dimensions without strict assumptions.
problem Understanding global convergence of ReLU nets in very high dimensions.
method Fine-grained analysis of random activation matrices and detailed gradient norm and curvature analysis.
result Empirical loss function has favorable geometrical properties in the overparameterized setting.
New algebraic rules for 5D shapes based on 3D cocycles.
problem Creating rules for 5D shapes.
method Using simplicial 3-cocycles to parameterize heptagon relations.
result Parameterized heptagon relations for 5D shapes.
The success of kernel-based learning methods depend on the choice of kernel. Recently, kernel learning methods have been proposed that use data to select the most appropriate kernel, usually by combining a set of base kernels. We introduce a new algorithm for kernel learning that combines a {\em continuous set of base …
Unified framework for learning quantum models from limited measurements.
problem Sample complexity and measurement shots in classical learning of quantum models.
method Unified learning framework considering probabilistic quantum measurements.
result Asymmetrical effects and interplay of sample size and measurement shots on learning performance.
Algorithm estimates human decision-making in high-dimensional states with finite-time guarantees.
problem Estimating optimal policies and measures of fit in dynamic decision models with high-dimensional state spaces.
method Single-loop estimation algorithm with stochastic gradient steps for reward maximization.
result Algorithm converges to a stationary solution with finite-time guarantees and approximates maximum likelihood sublinearly.
Knoop enhances variable selection with over-parameterization and knockoffs.
problem Challenges of variable selection in high-dimensional datasets.
method Generates knockoff variables, integrates them into an over-parameterized model, and uses anomaly-based significance tests.
result Superior performance in variable selection compared to existing methods.
Introduces a neural network-based method for efficient state and parameter estimation in complex systems.
problem Efficiently estimating state paths and parameters from noisy measurements in high-dimensional nonlinear systems.
method Bayesian Information Field Theory with neural network parameterization and optimization algorithms.
result Proposes a method to simplify and enrich state path parameterizations using neural networks, improving inference accuracy.
The aim of this survey is to give an overview on the geometry of Einstein maximal globally hyperbolic 2+1 spacetimes of arbitrary curvature, conatining a complete Cauchy surface of finite type. In particular a specialization to the finite type case of the canonicla Wick rotation-rescaling theory, previously developed b…
We introduce a method for constructing skills capable of solving tasks drawn from a distribution of parameterized reinforcement learning problems. The method draws example tasks from a distribution of interest and uses the corresponding learned policies to estimate the topology of the lower-dimensional piecewise-smooth…
Twisted SL2C local systems on surfaces of finite type appear often in geometry and physics. Most of them arise geometrically as local systems of charts for pleated hyperbolic structures. Bonahon and Thurston's "shear-bend coordinates" parameterize these local systems of charts. On a surface …
We consider four-dimensional gravity coupled to a non-linear sigma model whose scalar manifold is a non-compact geometrically finite surface Σ endowed with a Riemannian metric of constant negative curvature. When the space-time is an FLRW universe, such theories produce a very wide generalization of two-field α-att…
The space of Gaussian measures on a Euclidean space is geodesically convex in the L2-Wasserstein space. This space is a finite dimensional manifold since Gaussian measures are parameterized by means and covariance matrices. By restricting to the space of Gaussian measures inside the L2-Wasserstein space, we manag…
New parameterization tackles stochasticity in weather models.
problem Uncertainty in small-scale processes in weather models.
method Bayesian neural network with Hamiltonian Monte Carlo for uncertainty quantification and memory.
result Shows skillful forecasts and trustworthy uncertainty quantifications.
Study describes singularities of height functions on specific singular surfaces.
problem Analyzing singularities of height functions on singular surfaces.
method Using geometric language and blowing-ups, investigate singularities of height functions and dual surfaces.
result Characterized singularities of height functions and dual surfaces on specific singular surfaces.
The true distribution parameterizations of commonly used image datasets are inaccessible. Rather than designing metrics for feature spaces with unknown characteristics, we propose to measure GAN performance by evaluating on explicitly parameterized, synthetic data distributions. As a case study, we examine the performa…
Study describes singularities of distance squared functions on singular surfaces.
problem Characterizing singularities of distance squared functions on singular surfaces.
method Using smooth map-germs Sk, Bk, Ck, and F4 singularities, the study describes singularities via blowing-ups. result Characterization of singularities of wave-fronts and caustics of singular surfaces.
New method predicts state evolution for non-first-order algorithms on nonconvex problems.
problem Analyzing nonconvex optimization problems with random data.
method Developed a state evolution for a broader class of algorithms including first-order and saddle point updates.
result Established rigorous state evolution predictions and finite-sample guarantees for non-first-order methods.
We classify the 6-dimensional Lie algebras that can be endowed with an abelian complex structure and parameterize, on each of these algebras, the space of such structures up to holomorphic isomorphism.
Empirical study compares wide neural networks to kernel methods, resolving open questions.
problem Understanding the relationship between wide neural networks and kernel methods.
method Large-scale empirical study using various neural network architectures and kernel methods.
result Wide neural networks outperform fully-connected finite-width networks in some cases, but underperform convolutional finite-width networks.
Extends Lannes-Quillen theorem to all profinite groups.
problem Conjugacy separability of p-torsion elements and finite p-subgroups. method Developed a theory of products for families of discrete torsion modules.
result Proves a full version of the Lannes-Quillen theorem for all profinite groups.
Ensembles of neural networks improve training dynamics and performance.
problem Improving neural network performance through model size increase.
method Defining collegial ensembles (CE) as multiple independent models trained as a single model, and using theoretical results on NTK to optimize architecture search.
result CE dynamics simplify and scale favorably, resembling wide models, and can be efficiently implemented using group convolutions and block diagonal layers.
The paper studies implicit regularization in over-parameterized models for high-dimensional data.
problem Understanding implicit regularization in over-parameterized models for high-dimensional data.
method The paper designs regularization-free algorithms for the high-dimensional single index model and provides theoretical guarantees for the induced implicit regularization phenomenon.
result The proposed methods achieve minimax optimal statistical rates of convergence and outperform classical methods with explicit regularization.
The recently proposed option-critic architecture Bacon et al. provide a stochastic policy gradient approach to hierarchical reinforcement learning. Specifically, they provide a way to estimate the gradient of the expected discounted return with respect to parameters that define a finite number of temporally extended ac…
New bounds show BBVI's gradient variance matches SGD conditions, improving parameterization efficiency.
problem Understanding and improving the convergence of black-box variational inference (BBVI).
method Showed BBVI satisfies matching gradient variance bounds corresponding to the ABC condition for smooth and quadratically-growing log-likelihoods.
result Proven BBVI's gradient variance matches SGD conditions, with superior dimensional dependence for mean-field parameterization.
This note optimizes distributions using kernel mean embeddings with a new parameterization.
problem Optimizing distributions using kernel mean embeddings is challenging due to the difficulty of characterizing probability distribution vectors.
method Proposes a new parameterization of positive functions using kernel sums-of-squares to fit distributions in the MMD geometry.
result Distributions with kernel sum-of-squares densities are dense in the MMD geometry, allowing optimization in the finite-sample setting.
This paper proves SGD converges to global minimum for over-parameterized ReLU networks.
problem Theoretical understanding of implicit neural networks is limited.
method Gradient flow analysis of ReLU activated implicit neural networks.
result Randomly initialized gradient descent converges to global minimum at a linear rate for square loss function in over-parameterized ReLU networks.
New insights into bias and variance in over-parameterized models.
problem Understanding bias and variance in over-parameterized models.
method Analytic expressions derived from statistical physics for two minimal models.
result Over-parameterized models can overfit even in noiseless conditions.
Convex geometry explains optimal neural network parameters.
problem Understanding optimal parameters in over-parameterized neural networks.
method Convex geometry, extreme points, linear spline interpolation, kernel matrix, cutting-plane algorithm.
result Optimal network parameters can be characterized as interpretable closed-form formulas.
Geometric Occam's Razor shapes deep learning solutions.
problem Understanding the regularization in over-parameterized neural networks.
method Analyzing the geometric model complexity and Dirichlet energy in neural networks.
result Over-parameterized neural networks are implicitly regularized by geometric model complexity.
We establish the existence of an integer degree for the natural projection map from the space of parameterizations of asymptotically conical self-expanders to the space of parameterizations of the asymptotic cones when this map is proper. As an application we show that there is an open set in the space of cones in the …
This paper improves reinforcement learning efficiency for large-scale MDPs.
problem High sample complexity in tabular RL settings with large state and action spaces.
method Model-based approach and Q-learning with linearly parameterized features.
result Provably efficient learning with sample complexity bounds.
Study reveals benign overfitting in time series models with over-parameterization.
problem Analyzing over-parameterized linear models with dependent time-series data.
method Developed an estimator using interpolation and derived non-asymptotic risk bounds.
result Risk bound is influenced by the coherence of temporal covariance matrices at different time steps.
Learning suitable latent representations for observed, high-dimensional data is an important research topic underlying many recent advances in machine learning. While traditionally the Gaussian normal distribution has been the go-to latent parameterization, recently a variety of works have successfully proposed the use…
The study identifies spurious correlations in high-dimensional regression and quantifies their impact.
problem Spurious correlations in high-dimensional regression models.
method Statistical characterization of spurious correlations, quantifying their amount via ridge regularization.
result The value of regularization strength that minimizes test loss is in an interval where spurious correlations increase.
It has been noted in existing literature that over-parameterization in ReLU networks generally improves performance. While there could be several factors involved behind this, we prove some desirable theoretical properties at initialization which may be enjoyed by ReLU networks. Specifically, it is known that He initia…