We prove overfitting in minimal and random NNs, tempering the effect.
problem Overfitting in minimal and random neural networks.
method Analyzing binary weight fitting to noisy data, proving overfitting is tempered.
result The overfitting of minimal and random neural networks is tempered.
Minimum Description Length prevents overfitting in noisy data.
problem Learning from noisy data with overfitting risk.
method Minimum Description Length learning rule with tempered guarantees.
result Tempered agnostic finite sample learning guarantees and asymptotic behavior characterization.
New study finds many neural networks are not benignly overfitting.
problem Understanding the behavior of overfitting in neural networks.
method Exploring kernel ridge regression and deep neural networks to identify overfitting behaviors.
result Many interpolating methods, including neural networks, exhibit tempered overfitting rather than benign or catastrophic.
New findings on how neural networks generalize with varying dimensions.
problem Understanding how neural networks generalize with different dimensions and noise levels.
method Study of 2-layer ReLU NNs in a classification setting, proving transitions in overfitting types.
result The type of overfitting transitions from tempered to benign as input dimension increases.
New bounds for KRR condition number reveal overfitting phenomena.
problem Characterizing overfitting in KRR with varying kernel spectral decay.
method Derived new bounds for kernel matrices, enhanced test error bounds, and identified feature independence role.
result Identified tempered and catastrophic overfitting phenomena.
New framework uses tempered optimism to handle imperfect experts in online learning.
problem Challenges of implicit optimism in practical online learning environments.
method Introduces tempered optimism as a framework for online non-convex learning, modifies existing algorithms.
result Demonstrates tempered optimism as a fruitful paradigm for online non-convex learning.
New algorithm reduces overfitting in neural networks.
problem Overfitting in neural networks.
method Integrates SMC with SGHMC for mini-batch sampling.
result SMCSGHMC outperforms SGD and deep ensembles.
The paper examines how spike strengths and alignments affect overfitting in linear regression models.
problem The impact of spike strengths and alignments on overfitting in linear regression models.
method Characterization of generalization error through exact expressions and analysis of spike strengths, aspect ratio, and target alignment.
result Increasing spike strength can lead to catastrophic overfitting before benign overfitting, especially in well-specified aligned problems.
Study the cost of overfitting in noisy KRR models.
problem Cost of overfitting in noisy kernel ridge regression.
method An agnostic view of overfitting cost as a function of sample size for any target function, using Gaussian universality ansatz and task eigenstructure.
result Characterization of benign, tempered, and catastrophic overfitting.
New insights into Nadaraya-Watson interpolators show varied generalization behaviors.
problem Understanding generalization of interpolating predictors, especially in noisy data.
method Revisiting Nadaraya-Watson estimator with a single hyperparameter.
result Multiple overfitting behaviors exist, ranging from catastrophic to tempered.
In this paper we demonstrate that tempering Markov chain Monte Carlo samplers for Bayesian models by recursively subsampling observations without replacement can improve the performance of baseline samplers in terms of effective sample size per computation. We present two tempering by subsampling algorithms, subsampled…
Researchers study the geometric properties of a specific type of stable processes.
problem Understanding the information geometry of tempered stable processes.
method Derivation of α-divergence, Fisher information matrices, and α-connections.
result Obtained Fisher information matrices and α-connections for statistical manifolds.
Unified theory for kernel regression generalizes well under realistic assumptions.
problem Analyzing kernel regression under realistic conditions.
method Unified theory providing rigorous bounds for various settings.
result Self-regularization phenomenon in kernel matrices enables good generalization.
The paper connects tempering and entropic mirror descent for sampling.
problem Sampling from a target distribution with known unnormalized density.
method Establishes the connection between tempering SMC and entropic mirror descent, deriving convergence rates and geometric insights.
result Tempering SMC iterates correspond to entropic mirror descent on the reverse KL divergence, providing new optimization perspectives.
We investigate the class of tempered stable distributions and their associated processes. Our analysis of tempered stable distributions includes limit distributions, parameter estimation and the study of their densities. Regarding tempered stable processes, we deal with density transformations and compute their p-var…
This research improves binary classification by balancing overfitting and generalization with a novel Bayesian approach.
problem Improving binary classification models to avoid overfitting and generalize well.
method Introduces a PAC-Bayes type learning rule with a balancing parameter λ to balance training error and KL divergence to a prior.
result A choice of λ ensures uniformly vanishing excess loss, even in the agnostic case, by under-regularizing or over-regularizing appropriately.
We introduce a new distance metric for non-linear embeddings of Tempered Exponential Measures.
problem Non-linear embeddings of Tempered Exponential Measures (TEMs).
method Parameterization of finite discrete TEMs via Legendre functions, introducing tempered Hilbert co-simplex distance.
result Established a generalization of the Hilbert log cross-ratio simplex distance to a tempered Hilbert co-simplex distance.
A definition for elliptical tempered stable distribution, based on the characteristic function, have been explained which involve a unique spectral measure. This definition provides a framework for creating a connection between infinite divisible distribution, and particularly elliptical tempered stable distribution, w…
Geometric tempering fails for Langevin dynamics, proving convergence limits.
problem Proving convergence and limitations of geometric tempering for Langevin dynamics.
method Theoretical investigation of geometric tempering using Langevin dynamics.
result Geometric tempering can lead to exponential time convergence and poor functional inequalities.
Improved model-based estimation through tempered Bayes filter.
problem Improving predictive accuracy in partially-observable stochastic systems.
method Developed tempered Bayes filter combining likelihood and full posterior tempering.
result Tempered Bayes filter achieves improved predictive performance over the Bayes filter baseline.
New adaptive temperature selection improves parallel tempering efficiency.
problem Enhancing mixing in multi-modal distributions using parallel tempering.
method Adaptive temperature selection using policy gradient approach.
result Lower integrated autocorrelation times achieved compared to traditional methods.
Probabilistic Latent Semantic Analysis is a novel statistical technique for the analysis of two-mode and co-occurrence data, which has applications in information retrieval and filtering, natural language processing, machine learning from text, and in related areas. Compared to standard Latent Semantic Analysis which s…
Accumulated stock returns exhibit tempered skew t-distribution.
problem Analyzing the distribution of stock returns over multiple days.
method Employing a tempered skew t-distribution model.
result Tempered skew t-distribution fits the distribution of accumulated stock returns well.
We offer new formulas for European option pricing under tempered stable processes.
problem Pricing European options under tempered stable processes.
method Series expansions for tempered stable densities and European option prices.
result Our formulas are hyperparameter-free and competitive with traditional methods.
New financial models use tempered stable subordination for better correlation dynamics.
problem Building financial models with better correlation dynamics.
method Introducing tempered stable Sato subordinators and additive inhomogeneous processes.
result The new process has time-dependent correlation, improving fit for financial data.
Improves neural network performance by dynamically adjusting model weights based on source reliability.
problem Training neural networks on data from unreliable sources leads to poor performance.
method Dynamic re-weighting strategy using likelihood tempering to adjust model weights based on estimated source reliability.
result Significant improvement in model performance when trained on mixtures of reliable and unreliable data sources.
Polynomial mixing times for simulated tempering in mixture sampling problems.
problem Sampling from mixtures of log-concave distributions with location shifts.
method Conductance decomposition applied to an auxiliary Markov chain on an augmented space.
result First polynomial-time guarantee for simulated tempering with MALA.
Geometric tempering improves sampling from distributions, with exponential convergence rates.
problem Sampling from probability distributions using gradient flow dynamics.
method Geometric tempering of the target distribution in Wasserstein and Fisher-Rao gradient flows.
result Exponential convergence in continuous and discrete time for geometric tempering.
New method improves sampling from complex, multi-peaked distributions.
problem Sampling from high-dimensional, multimodal distributions using HMC.
method Combines tempered HMC with automatic tuning strategies.
result Demonstrates more effective scaling with dimension than adaptive methods.
This work tackles GAN training instability through parallel tempering.
problem Training instability and mode collapse in GANs.
method Introduces a parallel tempering framework to stabilize GAN training.
result Significantly reduces gradient variance and improves training efficiency.
The multivariate version of the Mixed Tempered Stable is proposed. It is a generalization of the Normal Variance Mean Mixtures. Characteristics of this new distribution and its capacity in fitting tails and capturing dependence structure between components are investigated. We discuss a random number generating procedu…
In this short note, we show how the parallel adaptive Wang-Landau (PAWL) algorithm of Bornn et al. (2013) can be used to automate and improve simulated tempering algorithms. While Wang-Landau and other stochastic approximation methods have frequently been applied within the simulated tempering framework, this note demo…
Improved lower bound for parallel tempering's mixing time.
problem Slow convergence and mixing in multimodal target distributions.
method Presented a new lower bound for the spectral gap of parallel tempering.
result Improved the best existing bound on spectral gap with polynomial dependence on parameters.
Develops a Monte Carlo algorithm for tempered stable process extrema.
problem Calculating the extrema of exponentially tempered Lévy processes.
method Novel Monte Carlo algorithm based on increments of the process.
result Geometrically fast convergence and optimal computational complexity.
New method estimates tempered stable Lévy models with high accuracy.
problem Estimating volatility and jump intensity of tempered stable Lévy processes.
method Iterative method combining Truncated Realized Quadratic Variations and small-time approximations.
result Method outperforms existing alternatives in various scenarios.
The paper uses FRFT to fit GTS distribution to asset returns.
problem Modeling asset returns with GTS distribution.
method Fractional Fourier Transform (FRFT) for fitting.
result GTS distribution fits SPY ETF and Bitcoin BTC returns.
Bayesian classification improves with explicit aleatoric uncertainty.
problem Lack of aleatoric uncertainty representation in Bayesian classification.
method Explicitly account for aleatoric uncertainty using a Dirichlet observation model.
result Explicit aleatoric uncertainty improves performance of Bayesian neural networks.
Characterizes Lévy-driven Ornstein-Uhlenbeck processes linked to tempered stable distributions.
problem Understanding Lévy-driven Ornstein-Uhlenbeck processes and their properties.
method Characterizes the Lévy triplet and deduces transition laws for finite variation Ornstein-Uhlenbeck processes associated with tempered stable distributions.
result Provides algorithms for generating skeleton of Ornstein-Uhlenbeck processes related to exponentially-modulated tempered stable laws.
New theorem improves spectral gap for sampling from mixture distributions.
problem Sampling from multimodal distributions with simulated tempering.
method Introduced a decomposition theorem for the restricted spectral gap of simulated tempering.
result Lower bound on the restricted spectral gap for mixture distributions.
New subgroup found in Lie groups with unusual properties.
problem Finding discrete subgroups with specific properties in Lie groups.
method Constructing a specific subgroup of a higher rank Lie group.
result Found a new subgroup that is dense, discrete, non-lattice, and non-tempered.
We define the spaces of Schwartz functions, tempered functions and tempered distributions on manifolds definable in polynomially bounded o-minimal structures. We show that all the classical properties that these spaces have in the Nash category, as first studied in Fokko du Cloux's work, also hold in this generalized s…
Paper introduces deterministic EM approximations for non-convex likelihood functions.
problem Deterministic approximations for the E-step of EM algorithm are lacking.
method Developed a theoretical framework for deterministic approximations, analyzed Riemann sums and tempered EM.
result Proved convergence guarantees for deterministic approximations and new non-trivial temperature profiles.
Enhances gradient-based discrete samplers with parallel tempering for multimodal distributions.
problem Local minima in high-dimensional, multimodal discrete distributions.
method Combines parallel tempering with discrete Langevin proposal, using Metropolis criterion for swaps.
result Significantly faster mixing and better sampling from complex distributions.
FlowVAT improves variational inference for multi-modal distributions.
problem Mode-seeking behavior and collapse in variational inference for complex posteriors.
method Conditional tempering approach for normalizing flow variational inference.
result FlowVAT outperforms traditional and adaptive annealing methods in multi-modal distributions, finding more modes and achieving better ELBO values.
The paper improves importance sampling and MCMC methods for complex distributions.
problem Improving sampling efficiency for distributions with atoms or heavy tails.
method Develops minimax optimal trial distributions and importance-tempered MCMC.
result Importance-tempered MCMC can be uniformly ergodic for certain distributions.
An infinite parallel tempering bouncy particle sampler improves sampling efficiency for multimodal distributions.
problem Sampling from complex posterior distributions with high accuracy and efficiency.
method Introduced an infinite parallel tempering bouncy particle sampler (BPS-PT) to accelerate convergence.
result Demonstrated improved sampling efficiency for multimodal distributions through numerical simulations.
Boosting with tempered exponential measures improves AdaBoost's convergence rate.
problem Improving the convergence rate of AdaBoost.
method Introducing tempered exponential measures (TEMs) to generalize AdaBoost's approach.
result t-AdaBoost achieves an improved convergence rate compared to AdaBoost, especially for t∈[0,1). The study examines denoising and noisy-input regression under distribution shift, revealing double descent behavior and insights for data augmentation.
problem Understanding denoising in machine learning, especially under noisy inputs and distribution shift.
method Theoretical analysis of supervised denoising and noisy-input regression, considering low-rank data and proportional regime.
result The test error exhibits double descent under general distribution shift, indicating that overfitting the noise can be benign, tempered, or catastrophic.