Log-linear models are the popular workhorses of analyzing contingency tables. A log-linear parameterization of an interaction model can be more expressive than a direct parameterization based on probabilities, leading to a powerful way of defining restrictions derived from marginal, conditional and context-specific ind…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Inference in log-linear models scales linearly in the size of output space in the worst-case. This is often a bottleneck in natural language processing and computer vision tasks when the output space is feasibly enumerable but very large. We propose a method to perform inference in log-linear models with sublinear amor…
Log-linear models are a family of probability distributions which capture relationships between variables. They have been proven useful in a wide variety of fields such as epidemiology, economics and sociology. The interest in using these models is that they are able to capture context-specific independencies, relation…
Algorithm optimizes quantized isotonic regression with log-linear time updates.
The paper integrates multiple Gaussian process predictions using Monte Carlo sampling.
A flexible nonparametric online changepoint detection algorithm for high-frequency data.
We compare various extensions of the Bradley-Terry model and a hierarchical Poisson log-linear model in terms of their performance in predicting the outcome of soccer matches (win, draw, or loss). The parameters of the Bradley-Terry extensions are estimated by maximizing the log-likelihood, or an appropriately penalize…
Changepoint detection is a central problem in time series and genomic data. For some applications, it is natural to impose constraints on the directions of changes. One example is ChIP-seq data, for which adding an up-down constraint improves peak detection accuracy, but makes the optimization problem more complicated.…
We study surfaces in Euclidean space that are minimal for a log-linear density , where are real numbers not all zero. We prove that if a surface is -minimal foliated by circles in parallel planes, then these planes are orthogonal to the vector and the surface must…
CANN models improve insurance claim count predictions using telematics data.
Two log-linear approximations speed up optimal transport for deep learning applications.
FEM improves attention mechanisms by applying value-driven log-linear tilts.
Paper explores duality in DPPs using embedding structure analysis.
A new method for efficient BNC parameter estimation outperforms HDP smoothing.
Hidden variables are ubiquitous in practical data analysis, and therefore modeling marginal densities and doing inference with the resulting models is an important problem in statistics, machine learning, and causal inference. Recently, a new type of graphical model, called the nested Markov model, was developed which …
A new probabilistic mixup framework improves deep learning generalization.
Privacy-preserving inference for clinical trials using differential privacy.
In this paper, we classify the class of constant weighted curvature curves in the plane with a log-linear density, or in other words, classify all traveling curved fronts with a constant forcing term in The classification gives some interesting phenomena and consequences including: the family of curves conv…
A major problem for the learning of Bayesian networks (BNs) is the exponential number of parameters needed for conditional probability tables. Recent research reduces this complexity by modeling local structure in the probability tables. We examine the use of log-linear local models. While log-linear models in this con…
Paper addresses online alignment of large language models under uncertain preference feedback.
We introduce a new embarrassingly parallel parameter learning algorithm for Markov random fields with untied parameters which is efficient for a large class of practical models. Our algorithm parallelizes naturally over cliques and, for graphs of bounded degree, its complexity is linear in the number of cliques. Unlike…
Improved speech recognition with language model integration in sequence-to-sequence models.
A novel method for learning DAGs from positive-valued data.
New method decomposes KL error using refined information and mode interactions.
This paper considers the problem of Byzantine fault tolerance in distributed linear regression in a multi-agent system. However, the proposed algorithms are given for a more general class of distributed optimization problems, of which distributed linear regression is a special case. The system comprises of a server and…
Estimates population size using capture-recapture designs with binary indicators.
Proposes statistical inference for dependency knowledge graphs from EHR data.
Learning the Markov network structure from data is a problem that has received considerable attention in machine learning, and in many other application fields. This work focuses on a particular approach for this purpose called independence-based learning. Such approach guarantees the learning of the correct structure …
Efficiently reduces rank of non-negative matrices with quadratic time complexity.
New algorithms explain Naive Bayes classifiers in polynomial time and delay.
McKernel introduces a framework to use kernel approximates in the mini-batch setting with Stochastic Gradient Descent (SGD) as an alternative to Deep Learning. Based on Random Kitchen Sinks [Rahimi and Recht 2007], we provide a C++ library for Large-scale Machine Learning. It contains a CPU optimized implementation of …
New methods combine model predictions to avoid linear mixtures' limitations.
Neural models learn continuous-time Markov chain transition rates from data.
Calculation of the log-normalizer is a major computational obstacle in applications of log-linear models with large output spaces. The problem of fast normalizer computation has therefore attracted significant attention in the theoretical and applied machine learning literature. In this paper, we analyze a recently pro…
A -translating soliton with density vector is a surface in Euclidean space whose mean curvature satisfies , where is the Gauss map. We classify all -translating solitons that are invariant by a one-parameter group of translations and a one-parameter group of rotat…
HGConv uses HRR to efficiently detect malware, outperforming existing methods.
Learning a regression function using censored or interval-valued output data is an important problem in fields such as genomics and medicine. The goal is to learn a real-valued prediction function, and the training output labels indicate an interval of possible values. Whereas most existing algorithms for this task are…
Log-linear models are arguably the most successful class of graphical models for large-scale applications because of their simplicity and tractability. Learning and inference with these models require calculating the partition function, which is a major bottleneck and intractable for large state spaces. Importance Samp…
We present a novel blind source separation (BSS) method, called information geometric blind source separation (IGBSS). Our formulation is based on the log-linear model equipped with a hierarchically structured sample space, which has theoretical guarantees to uniquely recover a set of source signals by minimizing the K…
New algorithm reconstructs sparse networks in subquadratic time.
We extend the theory of asymmetric information in mispricing models for stocks following geometric Brownian motion to constant relative risk averse investors. Mispricing follows a continuous mean--reverting Ornstein--Uhlenbeck process. Optimal portfolios and maximum expected log--linear utilities from terminal wealth f…
Reducing ICD-10 code granularity improves cost model accuracy and stability.
Standard autoregressive seq2seq models are easily trained by max-likelihood, but tend to show poor results under small-data conditions. We introduce a class of seq2seq models, GAMs (Global Autoregressive Models), which combine an autoregressive component with a log-linear component, allowing the use of global \textit{a…
AFTNet uses a network-constrained Weibull model for biomarker discovery.
Privacy preserving mechanisms such as differential privacy inject additional randomness in the form of noise in the data, beyond the sampling mechanism. Ignoring this additional noise can lead to inaccurate and invalid inferences. In this paper, we incorporate the privacy mechanism explicitly into the likelihood functi…
This thesis tackles Gaussian Process challenges in low dimensions.
In this paper we show that for the purposes of dimensionality reduction certain class of structured random matrices behave similarly to random Gaussian matrices. This class includes several matrices for which matrix-vector multiply can be computed in log-linear time, providing efficient dimensionality reduction of gene…
It is known that evolution strategies in continuous domains might not converge in the presence of noise. It is also known that, under mild assumptions, and using an increasing number of resamplings, one can mitigate the effect of additive noise and recover convergence. We show new sufficient conditions for the converge…