Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

11.5%23.1%34.6%46.2% · Jun 202019922001200920182026
48 results for data priors

This paper uses reference priors to improve deep learning models with unlabeled and labeled data.

problem Improving deep learning models with limited labeled data and unlabeled data from the same or related tasks.
method Develops and applies generalizations of reference priors for deep networks to exploit unlabeled and labeled data.
result Demonstrates new semi-supervised learning and pretraining methods for transfer learning.

Unified approach to estimating class-prior and classifier from PU data.

problem Learning binary classifier from positive and unlabeled data with class-prior estimation.
method Alternately estimating class-prior and training classifier.
result Unified approach improves classifier performance by accounting for class-prior estimation error.

Introduces foundation priors for using model-generated data in empirical research.

problem Using model-generated data as real observations in empirical research.
method Introduces foundation priors as an exponential-tilted, generalized Bayesian update of the user's primitive prior.
result Synthetic data reflects both model patterns and user's priors, enabling principled use in empirical work.

PriorGrad improves speech synthesis models by using data-dependent adaptive priors.

problem Inefficiency in denoising diffusion models due to mismatch between prior and data distributions.
method Proposes PriorGrad, an adaptive prior derived from data statistics based on conditional information.
result PriorGrad achieves faster convergence and superior performance in speech synthesis models.

Combines spatial context priors with data term for better dendritic spine segmentation.

problem Poor segmentation results due to overlapping pixel intensity distributions.
method Combines nonparametric context priors with learned-intensity data term and nonparametric shape priors.
result Significant improvements in dendritic spine segmentation.

A new data-adaptive prior stabilizes kernel learning in operators.

problem Learning kernels in operators from data is ill-posed due to nonlocal dependence.
method Introduces a data-adaptive prior to stabilize the Bayesian posterior mean.
result The data-adaptive prior achieves a stable posterior with small noise limits.

This paper proposes learning priors for adversarial autoencoders to improve model expressiveness.

problem The choice of priors in deep latent factor models can significantly affect model expressiveness, especially for models with limited capacity.
method The authors introduce code generators to transform simple priors into ones that better characterize the data distribution for adversarial autoencoders.
result The proposed model generates better image quality and learns better disentangled representations than standard AAEs in supervised and unsupervised settings.

Study proposes learning optimal priors from data for better Bayesian inference.

problem Challenges the use of noninformative uniform priors in Bayesian inference.
method Machine learning approach to learn optimal priors from data using a target function.
result Study models consistently outperformed baseline models in Wikipedia category classification.

Paper introduces a method to generate physically feasible dynamics with physical priors.

problem Challenges in generating physically feasible dynamics under physical priors.
method Seamlessly incorporates physical priors into diffusion-based generative models.
result Efficient generation of physically realistic dynamics across various physical phenomena.

TSFlow uses Gaussian processes to match priors for better time series forecasting.

problem Difficulties in aligning generative models' priors with time series data.
method Conditional flow matching (CFM) with Gaussian processes, optimal transport, and data-dependent priors.
result TSFlow produces high-quality unconditional samples and competitive forecasting results.

Data-dependent PAC-Bayes priors via differential privacy improve generalization bounds.

problem Creating valid generalization bounds for unknown data distributions.
method Using ε-differential privacy to construct data-dependent priors, leading to valid PAC-Bayes bounds.
result Data-dependent priors via differential privacy yield nonvacuous generalization bounds.

Improved Bayesian inference using power priors with historical data.

problem Improving Bayesian inference with historical data.
method Generalized power priors that adapt to the α\alpha parameter of Amari's α\alpha-divergence.
result Improved performance through appropriate choices of the α\alpha parameter.

X-VAE uses data-adaptive Gaussian priors to improve latent space modeling.

problem Limitations of standard Gaussian priors in complex datasets.
method Data-adaptive Gaussian prior derived from pretrained autoencoder latent codes.
result Improved latent space modeling and generation quality.

Improved VAE with optimal but intractable prior using density ratio trick.

problem Over-regularization with standard Gaussian prior in VAE.
method Introduced density ratio trick to estimate KL divergence without modeling aggregated posterior explicitly.
result VAE achieves high density estimation performance with implicit optimal prior.

A technique called 'prior laundering' uses legacy reconstructions to create uncertainty in Bayesian inverse problems.

problem Uncertainty in Bayesian inverse problems when data is uninformative.
method Using an archive of legacy reconstructions to create uncertainty in the posterior distribution, averaging the legacy posterior over measurements.
result The uncertainty reported in the posterior is inherited from the legacy reconstructions, not from the data itself.

This paper examines how the choice of prior distribution affects likelihoods of out-of-distribution inputs in deep generative models.

problem Mismatch between prior and data distributions causes deep generative models to assign higher likelihoods to out-of-distribution inputs.
method Proposes using a mixture distribution as a prior to make likelihoods of out-of-distribution inputs more sensitive.
result A mixture prior lowers the out-of-distribution likelihood with respect to real image data sets.

The MEM method uses data-driven priors for linear inverse problems, proving convergence and estimating differences.

problem Linear inverse problems with approximate priors.
method Maximum Entropy on the Mean (MEM) method with data-driven priors.
result Empirical mean convergence and estimates for prior differences based on epigraphical distance.

Proposes deep weight prior for improving neural network performance.

problem Improving neural network performance with limited training data.
method Defines deep weight prior (DWP) as an implicit distribution and proposes variational inference methods.
result Improves performance of Bayesian neural networks with limited data and accelerates conventional CNN training.

Unified analysis of privacy leakage in correlated data considering prior knowledge.

problem Understanding the impact of prior knowledge and data correlation on privacy leakage.
method Proposed prior differential privacy (PDP) and analyzed using WHG and multivariate Gaussian models.
result Derived closed-form expression for continuous data and chain rule for discrete data.

SIGMA prior enables federated learning for non-factorizable models.

problem Current FL methods assume conditional independence, limiting applicability to non-factorizable models.
method SIGMA prior approximates deep generative model to induce conditional independence structure.
result SIGMA prior expands FL applicability to fields requiring modeling dependencies.

Paper proposes a method to estimate class prior from positive and unlabeled data.

problem Estimating class prior in unlabeled datasets when labeled data is not available.
method Use penalized divergences to fit a mixture of class-wise distributions to the unlabeled data distribution.
result Correct estimation of class prior using only positive samples and penalized L1L_1-distance.

Maximizing mutual information selects simple models from limited data.

problem Selecting simple models from finite and potentially noisy data.
method Prior choice that maximizes mutual information between parameters and predictions.
result The method selects a lower-dimensional effective theory by ignoring poorly constrained parameters.

Estimates class prior for unlabeled data using kernel embedding.

problem Estimating class prior in PU learning scenario where only positive and full population samples are available.
method Direct estimator based on distribution matching and kernel embedding in Reproducing Kernel Hilbert Space.
result Asymptotic consistency and explicit deviation bound for the estimator.

Learning the network structure underlying data is an important problem in machine learning. This paper introduces a novel prior to study the inference of scale-free networks, which are widely used to model social and biological networks. The prior not only favors a desirable global node degree distribution, but also ta…

2015-03-07abs ↗pdf ↗

The paper proposes a method to integrate prior information into penalized regression.

problem Improving predictive performance in high-dimensional tasks with prior information.
method Integrating multiple sources of prior information into penalized regression.
result The method improves predictive performance, as shown by simulations and applications.

Improves latent space structure for better data representation.

problem Limited ability of conventional priors to encode data manifold structure.
method Introduces an Encoded Prior Sliced Wasserstein AutoEncoder with iterative training and geodesic interpolation.
result Learned manifold encoding preserves topological and geometric properties of data.

We use diffusion models to sample from complex GP priors in climate data.

problem Sampling from non-stationary Gaussian process priors is computationally hard.
method Replace GP prior with a diffusion model surrogate and use training-free guidance algorithms.
result Generated distributions are close to GP priors and can be fine-tuned.