The article derives a novel Gram-Charlier A (GCA) Series based Extended Rule-of-Thumb (ExROT) for bandwidth selection in Kernel Density Estimation (KDE). There are existing various bandwidth selection rules achieving minimization of the Asymptotic Mean Integrated Square Error (AMISE) between the estimated probability d…
Develops a neural network approach to solve inverse stochastic problems from particle observations.
problem Inference of Fokker-Planck equation coefficients from sparse particle data.
method Physics-informed neural networks (PINNs) with Kullback-Leibler divergence loss.
result Simultaneous inference of Fokker-Planck equation and multi-dimensional PDF from few particle observations.
New ICA algorithm improves source PDF estimation for better performance.
problem Inaccurate estimation of source PDFs leads to poor ICA performance.
method Entropy maximization with kernels, using global and local constraints.
result ICA-EMK outperforms competing algorithms in simulations and real-world data.
The Inverse Bagging Algorithm detects anomalies by identifying sub-samples rich in known data.
problem Detecting anomalies in data sets with a well-modeled process and an unknown PDF.
method Uses inverse bootstrap aggregating to identify sub-samples rich in the known process and classify events.
result The method avoids modifying the kinematic distributions of the well-modeled process.
Paper proposes using generalized lambda distributions for stochastic simulators.
problem Uncertainty quantification with complex stochastic models is computationally challenging.
method Flexible generalized lambda distribution approximates response PDF, parameters are sparse polynomial chaos expansions.
result Local inference of response PDF at each point of experimental design using replicated model evaluations.
Graph-based LRE estimates likelihood-ratios collaboratively for nodes.
problem Comparing unknown pdfs at graph nodes with graph structure.
method Graph-based Relative Unconstrained Least-squares Importance Fitting (GRULSIF).
result Collaborative estimation improves performance compared to independent methods.
Estimates class posterior probabilities without using scores from classifiers.
problem Estimating class posterior probabilities for new points in classification tasks.
method Varying prior probabilities to derive the ratio of pdf's at point x, directly determining class posterior probabilities.
result A method to estimate posterior probabilities without relying on classification scores.
Most signal processing problems involve the challenging task of multidimensional probability density function (PDF) estimation. In this work, we propose a solution to this problem by using a family of Rotation-based Iterative Gaussianization (RBIG) transforms. The general framework consists of the sequential applicatio…
New MC simulation methods use classifiers to estimate pdf ratios without explicit pdfs.
problem Estimating ratios of probability density functions (pdfs) without explicit pdfs.
method Proposes classifier-based pdf-free versions of MC simulation algorithms.
result Enables pdf-free simulation algorithms using surrogate functions computed by classifiers.
This paper improves parameter estimation in cardiac models using Gaussian process-based MH sampling.
problem Uncertainty in estimating patient-specific model parameters from sparse and noisy clinical data.
method Integrates surrogate modeling into Metropolis-Hastings sampling to improve computational efficiency and accuracy.
result Significant gain in computational efficiency without compromising accuracy, and insights into tissue heterogeneity.
Generative adversarial networks sample unknown high-dimensional conditional distributions.
problem Sampling from unknown high-dimensional conditional distributions with limited data.
method Generative adversarial networks (GAN) for both sampling and distribution inference.
result GAN effectively samples target conditional distribution with minimal impact on sample quality.
We investigate the historical volatility of the 100 most capitalized stocks traded in US equity markets. An empirical probability density function (pdf) of volatility is obtained and compared with the theoretical predictions of a lognormal model and of the Hull and White model. The lognormal model well describes the pd…
This research trains a supervised model to accurately detect PDF headings.
problem Detecting headings in PDFs for text extraction.
method Supervised learning with recursive feature elimination.
result Best classifier achieved 96.95% accuracy, 0.986 sensitivity, and 0.953 specificity.
Density destructors simplify complex PDFs to maximize entropy, linking to information theory.
problem Complex multivariate PDFs are hard to analyze.
method Invertible transforms that progressively remove structure from PDFs.
result Density destructors can improve estimates of information theoretic quantities.
New method calibrates photometric redshift PDFs more accurately.
problem Inaccurate photometric redshift uncertainties lead to systematic errors.
method Local re-calibration using feature-space regression of Probability Integral Transform (PIT) distributions.
result Calibrated PDFs are more accurate at all locations in feature space.
This work improves density estimation by characterizing pdf complexity using NL-spectrum.
problem Improving density estimation rates for general probability densities.
method Introducing NL-spectrum to characterize pdf complexity and deriving dimension-independent rates of convergence.
result Dimension-independent rates of convergence for fast density estimation.
SINF models transform arbitrary PDFs to target PDFs using 1D slices.
problem Transforming arbitrary probability distributions to target distributions efficiently.
method Iterative Optimal Transport of 1D slices, maximizing Wasserstein distance.
result SINF models generate high-quality samples and competitive density estimates.
We consider the problems of clustering, classification, and visualization of high-dimensional data when no straightforward Euclidean representation exists. Typically, these tasks are performed by first reducing the high-dimensional data to some lower dimensional Euclidean space, as many manifold learning methods have b…
Autoencoder optimizes data embedding for accurate PDF reproduction.
problem Inaccurate PDF reproduction in latent space of VAEs.
method Rate-Distortion Optimization guided autoencoder with isometric property.
result Our method achieves isometric data embedding and tractable PDF relations.
The paper introduces flat-topped PDFs for better fitting machine learning models.
problem Improving goodness of fit in machine learning models.
method Developed a new PDF based on the Fermi-Dirac or logistic function for adaptability.
result Flat-topped PDFs enhance model simplicity and fit quality.
An adaptive filter improves state estimation for complex systems.
problem Estimating non-Gaussian, multimodal PDFs in nonlinear systems.
method Adaptive split-combine Gaussian mixture filter (AMF) that splits and combines Gaussian particles adaptively.
result AMF consistently outperforms other filters across diverse benchmarks.
In the Black-Scholes context we consider the probability distribution function (PDF) of financial returns implied by volatility smile and we study the relation between the decay of its tails and the fitting parameters of the smile. We show that, considering a scaling law derived from data, it is possible to get a new f…
We report the proof that the expression of extended Gibrat's law is unique and the probability distribution function (pdf) is also uniquely derived from the law of detailed balance and the extended Gibrat's law. In the proof, two approximations are employed that the pdf of growth rate is described as tent-shaped expone…
CDF2PDF improves SIC for high-dimensional data estimation.
problem Estimating PDF from CDF in high-dimensional data.
method CDF2PDF approximates PDF by approximating CDF, avoiding hyper-parameter tuning and enabling polynomial time higher order derivative computation.
result CDF2PDF shows promising results in one-dimensional data experiments.
DeepPDF uses neural networks to estimate complex data distributions efficiently.
problem Efficiently estimating complex data distributions with high accuracy.
method DeepPDF uses a neural network to approximate a target pdf given samples, employing Probabilistic Surface Optimization (PSO) for stochastic optimization.
result DeepPDF achieves high inference accuracy for a wide range of target pdfs using a simple network structure.
Bayesian method estimates line frequencies with uncertainty.
problem Bayesian estimation of continuous frequencies.
method Variational Bayesian inference with von Mises mixtures.
result Significantly improved performance over point estimates.
We report the proof that the extension of Gibrat's law in the middle scale region is unique and the probability distribution function (pdf) is also uniquely derived from the extended Gibrat's law and the law of detailed balance. In the proof, two approximations are employed. The pdf of growth rate is described as tent-…
Bayesian method improves nanowire sensor parameter estimation.
problem Improving nanowire sensor parameter estimation.
method Bayesian inversion using PDE model and adaptive Metropolis algorithm.
result Simultaneous determination of nanowire sensor and analyte molecule properties.
Financial losses follow earthquake-like patterns, study finds.
problem Analyzing the timing between financial market losses.
method Fitting empirical interevent times with a Hawkes process.
result Financial market losses exhibit long-term memory similar to earthquakes.
In data science, it is often required to estimate dependencies between different data sources. These dependencies are typically calculated using Pearson's correlation, distance correlation, and/or mutual information. However, none of these measures satisfy all the Granger's axioms for an "ideal measure". One such ideal…
Deep learning reduces noise in weak lensing mass maps using GANs.
problem Noise reduction in weak lensing mass maps.
method Generative adversarial networks (GANs) applied to Subaru Hyper Suprime-Cam data.
result GANs successfully reproduce non-Gaussian information in denoised maps, showing stronger cosmological dependence.
Study fits BTC future returns from inverse options using logistic distribution.
problem Modeling future price distribution of Bitcoin.
method Fits empirical BTC future returns with logistic distribution using inverse options prices.
result BTC future returns can be described with a logistic distribution, but not stochastically.
DeepGDL models create realistic power grids from confidential data.
problem Creating realistic power grids from confidential data.
method Graph distribution learning (GDL) with a deep nonlinear recurrent structure.
result DeepGDL models accurately create synthetic power grids.
A step by step procedure to derive analytically the exact dynamical evolution equations of the probability density functions (PDF) of well known kinetic wealth exchange economic models is shown. This technique gives a dynamical insight into the evolution of the PDF, e.g., allowing the calculation of its relaxation time…
New framework quantifies uncertainty in data and models using RKHS.
problem Quantifying uncertainty in data and models.
method Projecting data into RKHS, transforming PDF, decomposing gradient flow.
result Decomposes uncertainty moments, providing discriminative resolution.
Physics-informed neural networks approximate diffusion process pdfs efficiently.
problem Approximating the probability density function of diffusion processes.
method Physics-informed neural networks solving Fokker-Planck or integro-differential equations.
result Neural network solutions approximate target solutions for various types of differential equations.
A new method uses histogram transform for better speaker identification.
problem Improving text-independent speaker identification.
method Uses Mel-frequency Cepstral coefficients and dynamic information among adjacent frames. Designs super-MFCCs features by cascading three neighboring MFCCs frames. Estimates PDF using histogram transform to generate more training data and reduce discontinuity.
result The histogram transform method shows improvement in speaker identification performance compared to conventional methods.
A new filter estimates complex system states more accurately.
problem Non-Gaussian features in nonlinear systems violate Kalman-type filters.
method Adaptive split-combine Gaussian mixture filter (AMF) that splits and combines Gaussian particles.
result AMF consistently outperforms other filters across diverse benchmarks.
A new classification method using class-specific features for improved text categorization.
problem Improving text categorization accuracy by leveraging class-specific features.
method EEF classifier based on class-specific features and optimal Bayesian classification rule.
result The proposed EEF classifier outperforms conventional methods on real-life data sets.
Many financial variables are found to exhibit multifractal nature, which is usually attributed to the influence of temporal correlations and fat-tailedness in the probability distribution (PDF). Based on the partition function approach of multifractal analysis, we show that there is a marked finite-size effect in the d…
A new visualization tool MD plot discovers interesting structures in continuous features.
problem Identifying interesting structures in data distributions, especially with skewed, clipped, or multimodal distributions.
method Proposes a new visualization tool called the mirrored density plot (MD plot) that does not require adjusting density estimation parameters.
result The MD plot outperforms conventional methods in identifying structures in complex distributions.
Proposes a deep learning method for uncertainty propagation in complex systems.
problem Uncertainty propagation in nonlinear dynamic systems with many uncertain variables.
method Data-driven approach using deep learning to approximate PDFs of uncertain systems.
result Demonstrates robustness evaluation of a feedback controller for a six-dimensional system.
I propose a frequency domain adaptation of the Expectation Maximization (EM) algorithm to group a family of time series in classes of similar dynamic structure. It does this by viewing the magnitude of the discrete Fourier transform (DFT) of each signal (or power spectrum) as a probability density/mass function (pdf/pm…
The paper integrates multiple Gaussian process predictions using Monte Carlo sampling.
problem Accurate prediction of variables using multiple models.
method Log-linear pooling of Gaussian process predictions, combined with Monte Carlo sampling.
result The log-linear pooling method improves prediction accuracy compared to linear pooling.
The Fisher information matrix (FIM) is a foundational concept in statistical signal processing. The FIM depends on the probability distribution, assumed to belong to a smooth parametric family. Traditional approaches to estimating the FIM require estimating the probability distribution function (PDF), or its parameters…
Unified framework for portfolio optimization using gain PDF.
problem Optimizing portfolios with control over high profits.
method Unified approach incorporating various PO methods using gain PDF.
result Directly matching target PDF for maximal control over PO.
Unified framework for PDF estimation using MDL-based binning and tensor factorization.
problem Challenges in estimating PDFs for non-uniform, multimodal data.
method MDL-based binning with quantile cuts, tensor factorization (CPD).
result Effective PDF estimation on synthetic and real data.
In this paper, a nonparametric maximum likelihood (ML) estimator for band-limited (BL) probability density functions (pdfs) is proposed. The BLML estimator is consistent and computationally efficient. To compute the BLML estimator, three approximate algorithms are presented: a binary quadratic programming (BQP) algorit…