The strong predictable representation property is proven for filtrations with a random variable under a density hypothesis.
problem Proving the strong predictable representation property in filtrations with a random variable.
method Using the density hypothesis of Jacod (1985), the strong predictable representation property is transferred to the enlarged filtration.
result The strong predictable representation property can always be transferred to the enlarged filtration under the density hypothesis.
New bounds on generalization error using information density moments.
problem Bounding the generalization error of randomized learning algorithms.
method Derives bounds on average and tail probabilities of generalization error using mth central moments of the information density.
result Explicit bounds on generalization error are derived, showing better dependence on confidence level with higher-order information density moments.
New method selects optimal bandwidth for price return density estimation, impacting efficient market hypothesis evaluation.
problem Estimating the complexity of price return distributions using kernel density estimation.
method Proposes a new complexity measure to select optimal bandwidth, avoiding overfitting and underfitting.
result Optimal bandwidth selection leads to clearer evaluation of the efficient market hypothesis.
Improved OOD detection using label smoothing and k-NN density estimates.
problem Detecting out-of-distribution examples in classification models.
method Label smoothing and k-NN density estimate on intermediate activations.
result Label smoothing improves OOD detection performance, both theoretically and empirically.
This paper tackles target-dependent label complexity gap in active learning.
problem Target-dependent label complexity gap in Agnostic Active Learning.
method Introduces a novel distribution-splitting strategy based on number density to reduce label complexity and error rate.
result Provides theoretical guarantees and practical advantages for reducing label complexity and error rate.
Study online monotone density estimation with expert aggregation and log-optimal calibration.
problem Online monotone density estimation and log-optimal calibration.
method Proposed two online estimators: Grenander estimator and expert aggregation estimator.
result Online estimators achieve O ( n 1 / 3 ) O(n^{1/3}) O ( n 1/3 ) cumulative log-likelihood gap and n log n \sqrt{n\log{n}} n log n pathwise regret bound. BMTI method estimates densities without bins, outperforming traditional estimators.
problem Nonparametric, robust, and data-efficient density estimation in high-dimensional spaces.
method BMTI integrates log-density differences between neighboring points, weighted by uncertainties, using a maximum-likelihood formulation.
result BMTI reconstructs smooth profiles in high-dimensional spaces, outperforming traditional estimators.
Paper finds wide minima are better for generalization and proposes a new learning rate schedule.
problem The challenge of finding optimal learning rates for model training.
method The paper introduces a new hypothesis about the density of wide minima and designs an explore-exploit learning rate schedule.
result The explore-exploit learning rate schedule improves model performance and reduces training time.
Study on hypothesis testing for densities and multinomials, showing local minimax rates and critical radii.
problem Testing goodness-of-fit for distributions with varying number of categories or unbounded support.
method Developed novel tests for both discrete and continuous cases, considering local minimax rates and critical radii.
result Characterized the dependence of critical radii on the null hypothesis and provided adaptive tests.
The need to estimate smooth probability distributions (a.k.a. probability densities) from finite sampled data is ubiquitous in science. Many approaches to this problem have been described, but none is yet regarded as providing a definitive solution. Maximum entropy estimation and Bayesian field theory are two such appr…
New bounds on learning algorithm generalization error derived using information density.
problem Bounding the generalization error of learning algorithms.
method Exponential inequalities and information density/conditional information density.
result Novel bounds on average and tail probability of generalization error.
Improved risk bounds for statistical inference problems.
problem Statistical inference and risk bounds for estimators.
method Adapted binary hypothesis testing approach to Fano's inequality for tighter lower bounds.
result Asymptotically sharp risk lower bounds for density estimation, active learning, and compressed sensing.
Paper proves diffusion models work on manifolds.
problem Current diffusion models assume densities are w.r.t. Lebesgue measure, limiting their applicability.
method Introduced convergence results for diffusion models on more general target distributions.
result Quantitative bounds on Wasserstein distance for target and generated distributions.
Diffusion models can generalize well even with coarse scores, thanks to the manifold hypothesis.
problem Understanding why diffusion models generate novel samples with coarse scores.
method Exploring the manifold hypothesis to explain diffusion model behavior.
result Diffusion models trained with coarse scores can achieve near-parametric rates of generalization, faster than estimating the full data distribution.
Power spectrum densities for the number of tick quotes per minute (market activity) on three currency markets (USD/JPY, EUR/USD, and JPY/EUR) for periods from January 1999 to December 2000 are analyzed. We find some peaks on the power spectrum densities at a few minutes. We develop the double-threshold agent model and …
The entropy density is an intuitive and powerful concept to study the complicated nonlinear processes derived from physical systems. We develop the minimum entropy density method (MEDM) to detect the structure scale of a given time series, which is defined as the scale in which the uncertainty is minimized, hence the p…
Unified Bayesian framework improves clinical trial hypothesis testing.
problem Lack of transparency and inability to quantify evidence in traditional P-values.
method Interval null hypothesis framework combined with Bayes factor-based tests.
result Bayesian interval hypothesis testing ensures frequentist error control and interpretability.
This paper resolves the test for Markov regime switching models' regime number.
problem Testing the number of regimes in Markov regime switching models.
method Derives the asymptotic distribution of the likelihood ratio test statistic.
result Establishes the asymptotic validity of the parametric bootstrap.
New test compares data to ergodic Markov models without specifying an alternative.
problem Testing goodness of fit for ergodic Markov processes without an alternative model.
method Density-based test comparing data to specified models' stationary densities.
result Test provides new insights into econometric and financial modeling.
Develops neural network methods for likelihood ratio estimation.
problem Estimating likelihood ratios from data.
method Neural network training for likelihood ratio estimation.
result Unified methodology for defining and solving optimization problems.
We examine the vertical component of surface area in the warped product of a Euclidean interval and a fiber manifold with product density. We determine general conditions under which vertical fibers minimize vertical surface area among regions bounding the same volume and use these results to conclude that in many such…
Develops a new test for comparing two groups' densities, showing minimax optimality.
problem Comparing probability densities between two groups.
method Probabilistic tensor product smoothing spline framework for joint density modeling; penalized likelihood ratio test for interaction testing.
result Proposed test is minimax optimal and outperforms conventional approaches.
A novel density-based approach QC detects outliers in data with high precision.
problem Detecting outliers in data with high precision and sensitivity.
method Quantum Clustering (QC) approach based on the density of data points.
result QC effectively finds hidden outliers and subtle outliers with parameter adjustment.
New algorithm selects best distribution privately in nearly-linear time.
problem Estimating the best distribution from samples under differential privacy constraints.
method Differentially private algorithm with nearly-linear time complexity and optimal approximation factor.
result Achieves optimal approximation factor of 3 with modest sample complexity increase.
Normal-bundle bootstrap generates new data preserving geometric structure.
problem Probabilistic models often exhibit salient geometric structure.
method NBB method decomposes probability measure into manifold and normal spaces, estimates manifold as density ridge, and generates new data by bootstrapping projection vectors.
result NBB generates new data that preserves the geometric structure of a given data set.
Unified view of score estimators for flexible densities.
problem Estimating the score from unknown distributions.
method Regularized nonparametric regression framework.
result Unified convergence analysis and new estimators with desirable properties.
The paper proposes a method to estimate joint probability from unpaired data using entropic transport kernels.
problem Estimating joint probability from unpaired data with unknown internal ordering.
method Maximum-likelihood inference, entropic optimal transport kernels, EMML algorithm.
result The method can recover true density from empirical approximations as the number of blocks increases.
Diffusion models adapt to data geometry through log-domain smoothing.
problem Understanding why diffusion models generalize well across diverse domains.
method Investigating the role of score matching and log-domain smoothing in diffusion models.
result Log-domain smoothing adapts the diffusion model to the data manifold.
Paper proposes methods to learn sub-manifolds and estimate densities in normalizing flows.
problem Normalizing flows struggle with finding sub-manifolds in high-dimensional data.
method Introduces per-pixel penalized log-likelihood and hierarchical training approaches.
result Validated superior performance in manifold learning and density estimation.
VAE with noise model learns smoothed densities without seeing noisy data.
problem Learning smoothed densities with noisy data.
method Imaginary noise model in variational autoencoders (σ-VAE).
result All σ-VAEs are equivalent via β-VAE expansion.
Optimal score function estimation via empirical risk minimization
problem Estimating the score function of a probability measure on the flat torus from a sample
method Constraining the hypothesis space to a Sobolev ball
result Minimax estimation rates are achieved
The paper tackles manifold overfitting in deep generative models.
problem Manifold overfitting occurs when generative models learn the manifold itself instead of the distribution on it.
method The authors propose a two-step procedure: dimensionality reduction followed by maximum-likelihood density estimation.
result The two-step procedure avoids manifold overfitting and enables density estimation on learned manifolds.
We propose a method to infer causal structures containing both discrete and continuous variables. The idea is to select causal hypotheses for which the conditional density of every variable, given its causes, becomes smooth. We define a family of smooth densities and conditional densities by second order exponential mo…
GANs learn from generated data without likelihoods.
problem Learning from implicit generative models without explicit likelihoods.
method Density ratio estimation and hypothesis testing.
result Derivation of GAN's objective function and related objectives.
Paper compares optimal denoising methods for generative models, finding different results based on data regularity.
problem Optimizing denoising in score-based generative models for various data types.
method Comparison of full-denoising and half-denoising approaches, analyzing performance in terms of distribution distances.
result Different denoising methods perform better under different data regularity conditions.
Paper proposes kernel-based tests for model misspecification.
problem Determining if a model is misspecified.
method Minimum distance estimators based on MMD and KSD.
result Correct test level maintained without data splitting.
New test detects conditional dependence using GAN approximations.
problem Detecting conditional dependence in high-dimensional feature spaces.
method Generative adversarial networks (GANs) for approximating conditional distributions.
result Significant gains in power over competing methods.
Study bounds on cusp volumes of alternating knots on surfaces.
problem Bounding cusp volumes of knots on surfaces.
method Analyzing hyperbolic knots with alternating projections on embedded surfaces.
result Two-sided bounds on cusp area in terms of twist number and surface genus.
Detects which features have shifted in data distributions.
problem Identifying which specific features have caused a distribution shift.
method Formalizes the problem as multiple conditional distribution hypothesis tests, proposes non-parametric and parametric statistical tests, and uses a test statistic based on the density model score function.
result Demonstrates methods for identifying when and where a shift occurs in multivariate time-series data.
New research determines the optimal sample complexity for multiclass and list learning.
problem Determining the optimal sample complexity for multiclass classification.
method Algebraic characterization of multiclass hypothesis classes in terms of their DS dimension.
result Proves a longstanding conjecture and determines the optimal dependence of sample complexity on DS dimension.
New approach avoids restrictive assumptions for optimal portfolio in default risk scenarios.
problem Optimal portfolio optimization under default risk when traditional techniques are not applicable.
method Alternative approach using forward integration to avoid Jacod density hypothesis.
result Weaker intensity hypothesis is the appropriate condition for optimality in logarithmic utility.
Study on sequential prediction with log-loss, focusing on well-specified and misspecified cases.
problem Sequential prediction with log-loss under different specification conditions.
method Analysis of cumulative regret in well-specified and misspecified cases for a Gaussian location hypothesis class.
result Cumulative regrets in well-specified and misspecified cases asymptotically coincide for the d d d -dimensional Gaussian location hypothesis class. Method calculates financial distributions using recursive relationships.
problem Analyzing the distribution of financial functions at discrete points.
method Recursive method to calculate probability distributions.
result High accuracy demonstrated in numerical experiments.
Paper uses machine learning to estimate IRI from pavement distress types, densities, and severities.
problem Costly IRI measurements exclude many road classes; estimating IRI from distress data is needed.
method Data from in-service pavements; machine learning methods used to predict IRI.
result Machine learning can reliably estimate IRI based on distress types, densities, and severities.
A new HMM model captures kernel dependencies using context-specific Bayesian networks.
problem Traditional HMMs struggle with non-Gaussian data and independence assumptions.
method Kernel density estimation with context-specific Bayesian networks.
result The proposed model outperforms related HMMs in likelihood and classification accuracy.
We identify a condition for regularity of optimal transport maps that requires only three derivatives of the cost function, for measures given by densities that are only bounded above and below. This new condition is equivalent to the weak Ma-Trudinger-Wang condition when the cost is C 4 C^4 C 4 . Moreover, we only require (n…
Joint sensing and communication network improves target localization and reduces communication load.
problem Efficiently localize multiple targets with reduced communication overhead.
method Multi-base station cooperative sensing with AI-aided clustering and tracking.
result Optimal sub-pattern assignment (OSPA) error less than 60 cm with reduced communication capacity.
A test for weak signal detection in noisy data matrices.
problem Detecting a weak signal in a noisy Wigner matrix when the signal-to-noise ratio is small.
method Utilizes linear spectral statistics and hypothesis testing on the data matrix.
result The proposed test is optimal when the noise is Gaussian and can be improved with known noise density.