Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

132263395526 · Jun 202019922001200920182026
48 results for Maximum Pseudolikelihood Estimation

I propose a variational approach to maximum pseudolikelihood inference of the Ising model. The variational algorithm is more computationally efficient, and does a better job predicting out-of-sample correlations than L2L_2 regularized maximum pseudolikelihood inference as well as mean field and isolated spin pair appro…

2014-09-24abs ↗pdf ↗

Paper proposes a new method for MRF structure learning.

problem Learning MRF structures without strong distributional assumptions.
method Grow-Shrink Maximum Pseudolikelihood Estimation (MPLE) for edge set optimization.
result The method successfully handles symmetricity in MRFs and outperforms previous methods.

Maximum pseudolikelihood method has been among the most important methods for learning parameters of statistical physics models, such as Ising models. In this paper, we study how pseudolikelihood can be derived for learning parameters of a mixture of Ising models. The performance of the proposed approach is demonstrate…

2015-06-08abs ↗pdf ↗

Two methods improve tensor recovery in Ising models, revealing gene interactions.

problem Improving tensor recovery in Ising models for complex data structures.
method Pseudolikelihood and interaction screening approaches for tensor learning.
result Both methods achieve tensor recovery with sample size logarithmic in nodes, exponential in strength and degree.

We consider the high-dimensional heteroscedastic regression model, where the mean and the log variance are modeled as a linear combination of input variables. Existing literature on high-dimensional linear regres- sion models has largely ignored non-constant error variances, even though they commonly occur in a variety…

2012-05-21abs ↗pdf ↗

Curating labeled training data has become the primary bottleneck in machine learning. Recent frameworks address this bottleneck with generative models to synthesize labels at scale from weak supervision sources. The generative model's dependency structure directly affects the quality of the estimated labels, but select…

2017-03-02abs ↗pdf ↗

Score matching efficiency tied to distribution isoperimetric properties.

problem Understanding when score matching is as efficient as maximum likelihood.
method Connecting score matching efficiency to isoperimetric constants of distributions.
result Score matching is statistically efficient when the distribution has a small isoperimetric constant.

Researchers use statistical methods to infer transmission matrices in complex media.

problem Comprehending and exploiting photon scattering through disordered media.
method Pseudolikelihood decimation to learn the coupling matrix via random sampling.
result Transmission matrices can be inferred and used like normal optical elements.

Study investigates learning performance in inverse Ising problems with sparse teacher couplings.

problem Learning performance in inverse Ising problems with sparse teacher couplings.
method Pseudolikelihood maximization method, replica and cavity methods from statistical mechanics.
result Perfect inference of teacher's couplings is possible in the thermodynamic limit for certain conditions.

Kernel methods are one of the mainstays of machine learning, but the problem of kernel learning remains challenging, with only a few heuristics and very little theory. This is of particular importance in methods based on estimation of kernel mean embeddings of probability measures. For characteristic kernels, which inc…

2016-03-07abs ↗pdf ↗

We consider a problem of data integration. Consider determining which genes affect a disease. The genes, which we call predictor objects, can be measured in different experiments on the same individual. We address the question of finding which genes are predictors of disease by any of the experiments. Our formulation i…

2016-10-03abs ↗pdf ↗

Estimates inverse temperature of Ising models with a single sample.

problem Estimating inverse temperature in truncated Ising models with hard constraints.
method Maximizing pseudolikelihood to estimate the inverse temperature.
result An estimator that is nearly O(n)O(n) time and O(Δ3/n)O(Δ^3/\sqrt{n})-consistent.

PPL improves on Takacs-Fiksel estimation for Gibbs models.

problem Improving point process estimation methods.
method PPL uses cross-validation and a specific loss function to estimate parameters.
result PPL with specific loss functions and hyperparameters outperforms Takacs-Fiksel estimation in mean square error.

New method identifies network structure without regularization for sparse teacher couplings.

problem Identifying network structure in inverse Ising problems with model mismatch.
method Ridge linear regression with two-stage estimator.
result Perfect identification of network structure possible without regularization for sparse teacher couplings.

Due to the intractable partition function, the exact likelihood function for a Markov random field (MRF), in many situations, can only be approximated. Major approximation approaches include pseudolikelihood and Laplace approximation. In this paper, we propose a novel way of approximating the likelihood function throug…

2018-03-27abs ↗pdf ↗

Trans-Ising combines auxiliary datasets to estimate high-dimensional Ising models.

problem Limited target sample sizes and difficulty in using auxiliary binary datasets of unknown relevance.
method Trans-Ising uses a loss-based source screening rule and a two-stage estimation procedure.
result Trans-Ising achieves lower estimation errors than target-only estimation and naive data pooling.

Two scalable methods for PSL structure learning improve runtime and AUC.

problem Efficiently learning clauses for probabilistic soft logic models.
method Greedy search and a novel optimization method combining data-driven clause generation and PPLL objective.
result PPLL achieves up to 15% AUC gains and an order of magnitude runtime speedup.

Estimates complex dependency structures in multi-omics data.

problem Graphical model estimation from multi-omics data with scalability and consistency.
method Pseudolikelihood-based graphical model framework with 1\ell_1-penalized empirical risk.
result Estimates partial correlation network from dual-omic liver cancer data.

FAST-DAD distills complex ensemble models into faster, more accurate individual models.

problem Deploying complex AutoML ensemble predictors on tabular data is slow, large, and opaque.
method Data augmentation strategy based on Gibbs sampling from a self-attention pseudolikelihood estimator.
result FAST-DAD distillation produces significantly better individual models than standard training.

Maximum likelihood estimation fails to be well-posed in Gaussian process regression.

problem Establishing well-posedness of maximum likelihood estimation in Gaussian process regression.
method Analyzing the conditions under which maximum likelihood estimation is not Lipschitz in the data with respect to the Hellinger distance.
result Maximum likelihood estimation is not well-posed in the noiseless data setting for any Gaussian process with a stationary covariance function whose lengthscale parameter is estimated using maximum likelihood.

We present a new statistical learning paradigm for Boltzmann machines based on a new inference principle we have proposed: the latent maximum entropy principle (LME). LME is different both from Jaynes maximum entropy principle and from standard maximum likelihood estimation.We demonstrate the LME principle BY deriving …

2012-10-19abs ↗pdf ↗

Study connects spectral clustering to maximum margin and level set estimation.

problem Connecting spectral clustering to maximum margin and level set estimation.
method Obtained bounds on eigenvectors of graph Laplacian matrices in terms of cluster separation and connectivity. Showed sensitivity mitigation by removing outliers and estimating level sets.
result Spectral clustering converges to maximum margin clustering as scaling parameter approaches zero.

We propose a robust estimator to improve maximum likelihood in probabilistic models.

problem Overfitting and sensitivity to noise in maximum likelihood estimation.
method Distributionally robust maximum likelihood estimator that minimizes worst-case expected log-loss.
result The robust estimator is statistically consistent and performs well in regression and classification tasks.

Paper presents a method for estimating Hawkes process parameters.

problem Estimating parameters of Hawkes processes with self-excitation or inhibition.
method Maximum likelihood estimation for Hawkes processes with self-excitation or inhibition.
result The proposed estimator provides more accurate estimations in the inhibition context.

For many large undirected models that arise in real-world applications, exact maximumlikelihood training is intractable, because it requires computing marginal distributions of the model. Conditional training is even more difficult, because the partition function depends not only on the parameters, but also on the obse…

2012-07-04abs ↗pdf ↗

A new method improves text generation quality and diversity.

problem Exposure bias in Maximum Likelihood Estimation for text generation.
method ψ-MLE, a new training scheme based on density ratio estimation.
result ψ-MLE outperforms Maximum Likelihood Estimation and other models in text generation quality and diversity.

We introduce a new method for training deep Boltzmann machines jointly. Prior methods of training DBMs require an initial learning pass that trains the model greedily, one layer at a time, or do not perform well on classification tasks. In our approach, we train all layers of the DBM simultaneously, using a novel train…

2013-01-16abs ↗pdf ↗

This paper addresses the estimation of parameters of a Bayesian network from incomplete data. The task is usually tackled by running the Expectation-Maximization (EM) algorithm several times in order to obtain a high log-likelihood estimate. We argue that choosing the maximum log-likelihood estimate (as well as the max…

2011-10-12abs ↗pdf ↗

We develop a maximum penalized quasi-likelihood estimator for estimating in a nonparametric way the diffusion function of a diffusion process, as an alternative to more traditional kernel-based estimators. After developing a numerical scheme for computing the maximizer of the penalized maximum quasi-likelihood function…

2010-08-14abs ↗pdf ↗

Investigates numerical issues in GP interpolation parameter estimation.

problem Numerical issues in maximum likelihood parameter estimation for Gaussian process interpolation.
method Investigates and proposes strategies to improve open-source software implementations.
result Improves reliability and reproducibility of studies relying on GP implementations.

A boosting method improves nonparametric density estimation without smoothing assumptions.

problem Overfitting in nonparametric data fitting.
method Introduces a boosting algorithm for univariate nonparametric maximum likelihood estimation.
result Demonstrates the effectiveness of the boosting approach through simulations and real data experiments.