Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

3867711,1571,542 · Jun 202019922001200920182026
48 results for Log-Linear Learning

Efficiently compares independence structures in log-linear models.

problem Limited direct measures for comparing log-linear model independence structures.
method Direct comparison method based on independence structure, efficient computation.
result First metric for direct comparison of log-linear model independence structures.

Log-linear models are the popular workhorses of analyzing contingency tables. A log-linear parameterization of an interaction model can be more expressive than a direct parameterization based on probabilities, leading to a powerful way of defining restrictions derived from marginal, conditional and context-specific ind…

2014-09-09abs ↗pdf ↗

Two log-linear approximations speed up optimal transport for deep learning applications.

problem Computing optimal transport in high dimensions is computationally expensive.
method Locality-sensitive hashing (LSH) and Nyström approximation with LSH-based sparse corrections.
result Log-linear time algorithms for entropy-regularized OT perform well in high-dimensional spaces.

GAMs combine autoregressive and log-linear components for data-efficient sequence learning.

problem Poor performance of standard autoregressive models under small-data conditions.
method Introduce Global Autoregressive Models (GAMs) combining autoregressive and log-linear components, trained in two steps.
result GAMs show a strong perplexity reduction over standard models in language modelling.

Models predict soccer match outcomes with similar accuracy.

problem Predicting soccer match outcomes (win, draw, loss).
method Compared Bradley-Terry extensions and hierarchical Poisson log-linear model. Parameters estimated using log-likelihood or integrated nested Laplace approximations. Predictive performance assessed using temporal validation.
result Bradley-Terry extensions and hierarchical Poisson log-linear model perform similarly in predicting match outcomes.

A major problem for the learning of Bayesian networks (BNs) is the exponential number of parameters needed for conditional probability tables. Recent research reduces this complexity by modeling local structure in the probability tables. We examine the use of log-linear local models. While log-linear models in this con…

2013-01-23abs ↗pdf ↗

We study surfaces in Euclidean space R3{\mathbb R}^3 that are minimal for a log-linear density φ(x,y,z)=αx+βy+γyφ(x,y,z)=αx+βy+γy, where α,β,γα,β,γ are real numbers not all zero. We prove that if a surface is φφ-minimal foliated by circles in parallel planes, then these planes are orthogonal to the vector (α,β,γ)(α,β,γ) and the surface must…

2014-10-09abs ↗pdf ↗

The paper integrates multiple Gaussian process predictions using Monte Carlo sampling.

problem Accurate prediction of variables using multiple models.
method Log-linear pooling of Gaussian process predictions, combined with Monte Carlo sampling.
result The log-linear pooling method improves prediction accuracy compared to linear pooling.

Paper explores duality in DPPs using embedding structure analysis.

problem Understanding the geometric structure of determinantal point processes.
method Analyzes the exponential family embedding of DPPs and uses the e-embedding curvature tensor.
result Discovers the duality between marginal and L-ensemble kernels.

Paper develops DLTF to learn optimized dictionaries for efficient thresholded feature recovery.

problem Efficiently recover sparse code support from time-consuming sparse coding.
method Formulates DLTF model to learn optimized dictionary for thresholded feature, derives log-linear time proximal operator.
result DLTF model demonstrates remarkable efficiency, effectiveness, and robustness in various tasks.

Hidden variables are ubiquitous in practical data analysis, and therefore modeling marginal densities and doing inference with the resulting models is an important problem in statistics, machine learning, and causal inference. Recently, a new type of graphical model, called the nested Markov model, was developed which …

2013-09-26abs ↗pdf ↗

New method decomposes KL error using refined information and mode interactions.

problem Learning probability distributions over discrete variables with higher-order interactions.
method Using information geometry, refined mode interactions, and a novel Monte-Carlo sampling technique.
result Complete decomposition of KL error and efficient data use.

A new method for efficient BNC parameter estimation outperforms HDP smoothing.

problem Efficiently estimating parameters for Bayesian network classifiers to match or exceed random forest performance.
method Uses log-linear regression to approximate hierarchical Dirichlet process (HDP) smoothing, making the approach simpler and faster.
result Our method outperforms HDP smoothing while being orders of magnitude faster and competitive with random forests.

HGConv uses HRR to efficiently detect malware, outperforming existing methods.

problem Efficiently detecting malware with long sequences.
method Holographic Global Convolutional Networks (HGConv) utilizing Holographic Reduced Representations (HRR).
result Achieved state-of-the-art results on malware benchmarks.

New methods combine model predictions to avoid linear mixtures' limitations.

problem Combining predictions from different models to avoid linear mixtures' limitations.
method Log-linear pooling (locking) and quantum superposition (quacking) to optimise model weights.
result Demonstrated locking method with illustrative example and practical application.

Learning the Markov network structure from data is a problem that has received considerable attention in machine learning, and in many other application fields. This work focuses on a particular approach for this purpose called independence-based learning. Such approach guarantees the learning of the correct structure …

2013-07-15abs ↗pdf ↗

Paper addresses online alignment of large language models under uncertain preference feedback.

problem Online alignment of large language models with misspecified preference feedback.
method Formulates an oracle-robust objective as a worst-case optimization problem for log-linear policies, and develops projected stochastic composite updates.
result Shows that the robust objective admits an exact closed-form decomposition and achieves O~(ε2)\widetilde{O}(\varepsilon^{-2}) oracle complexity.

Improved speech recognition with language model integration in sequence-to-sequence models.

problem Improving word error rate in speech recognition models.
method Log-linear combination of acoustic and language models with per-token renormalization.
result The proposed method shows good improvements over standard model combination on Librispeech system.

Calculation of the log-normalizer is a major computational obstacle in applications of log-linear models with large output spaces. The problem of fast normalizer computation has therefore attracted significant attention in the theoretical and applied machine learning literature. In this paper, we analyze a recently pro…

2015-06-12abs ↗pdf ↗

A flexible nonparametric online changepoint detection algorithm for high-frequency data.

problem Detecting changes in real-time in high-frequency data streams with limited computational resources.
method NP-FOCuS, a sequential likelihood ratio test for a change in the empirical cumulative density function, using functional pruning.
result NP-FOCuS outperforms current nonparametric online changepoint techniques in various settings.

We introduce a new embarrassingly parallel parameter learning algorithm for Markov random fields with untied parameters which is efficient for a large class of practical models. Our algorithm parallelizes naturally over cliques and, for graphs of bounded degree, its complexity is linear in the number of cliques. Unlike…

2013-08-29abs ↗pdf ↗

A novel method for learning DAGs from positive-valued data.

problem Causal discovery from observational data of positive-valued variables.
method Hybrid Moment-Ratio Scoring (H-MRS) algorithm combining moment-based scoring and log-scale regression.
result H-MRS integrates log-scale Ridge regression for moment-ratio estimation with a greedy ordering procedure based on raw-scale moment ratios, followed by Elastic Net-based parent selection.

CANN models improve insurance claim count predictions using telematics data.

problem Improving insurance claim count predictions with telematics data.
method Combining classical actuarial models with neural networks for telematics data.
result CANN models outperform traditional models in predicting insurance claims.

Efficiently reduces rank of non-negative matrices with quadratic time complexity.

problem Efficiently reducing the rank of non-negative matrices.
method Formulated rank reduction as a mean-field approximation using a log-linear model.
result Optimal solution for minimizing KL divergence can be computed in closed form.

Changepoint detection is a central problem in time series and genomic data. For some applications, it is natural to impose constraints on the directions of changes. One example is ChIP-seq data, for which adding an up-down constraint improves peak detection accuracy, but makes the optimization problem more complicated.…

2017-03-09abs ↗pdf ↗

New framework compares two stochastic learning dynamics in games.

problem Inability to distinguish between different learning rules leading to the same steady-state behavior.
method Developed a framework for comparative analysis of stochastic learning dynamics with different update rules.
result Identified distinct behaviors in the paths to stochastically stable states for LLL and ML.

Discriminative linear models are a popular tool in machine learning. These can be generally divided into two types: The first is linear classifiers, such as support vector machines, which are well studied and provide state-of-the-art results. One shortcoming of these models is that their output (known as the 'margin') …

2012-06-27abs ↗pdf ↗

Study on bias-variance trade-off in hierarchical models with higher-order interactions.

problem Understanding the bias-variance trade-off in hierarchical probabilistic models with higher-order interactions.
method Proposed an efficient inference algorithm using Gibbs sampling and annealed importance sampling for log-linear higher-order Boltzmann machine.
result Higher-order interactions produce less variance for smaller sample size and comparable error with hidden layers.

Estimates population size using capture-recapture designs with binary indicators.

problem Estimating population size from capture-recapture data with binary indicators.
method Proposes a modern method using undersmoothed lasso model to estimate the target parameter of interest.
result The choice of constraint on the K-dimensional distribution significantly impacts the value of the estimand.

Proposes a neural network for EMG classification that handles skewness and kurtosis.

problem EMG classification methods struggle with skewness and kurtosis in EMG feature distributions.
method Introduces a neural network based on the Johnson SUS_\mathrm{U} translation system.
result Achieves high classification performance without hyperparameter tuning.

We extend the theory of asymmetric information in mispricing models for stocks following geometric Brownian motion to constant relative risk averse investors. Mispricing follows a continuous mean--reverting Ornstein--Uhlenbeck process. Optimal portfolios and maximum expected log--linear utilities from terminal wealth f…

2011-01-06abs ↗pdf ↗

We study the problem of collaborative filtering where ranking information is available. Focusing on the core of the collaborative ranking process, the user and their community, we propose new models for representation of the underlying permutations and prediction of ranks. The first approach is based on the assumption …

2014-07-23abs ↗pdf ↗

Reducing ICD-10 code granularity improves cost model accuracy and stability.

problem High-dimensional regression with ICD-10 codes leads to unstable coefficient estimates.
method Log-linear analytics approach to cost model regularization through diagnostic code merging.
result Reducing ICD-10 code granularity from 7 characters to 6 or fewer improves model interpretability and consistency.