Algorithm optimizes quantized isotonic regression with log-linear time updates.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Inference in log-linear models scales linearly in the size of output space in the worst-case. This is often a bottleneck in natural language processing and computer vision tasks when the output space is feasibly enumerable but very large. We propose a method to perform inference in log-linear models with sublinear amor…
Log-linear models are the popular workhorses of analyzing contingency tables. A log-linear parameterization of an interaction model can be more expressive than a direct parameterization based on probabilities, leading to a powerful way of defining restrictions derived from marginal, conditional and context-specific ind…
Two log-linear approximations speed up optimal transport for deep learning applications.
New algorithms explain Naive Bayes classifiers in polynomial time and delay.
FEM improves attention mechanisms by applying value-driven log-linear tilts.
Changepoint detection is a central problem in time series and genomic data. For some applications, it is natural to impose constraints on the directions of changes. One example is ChIP-seq data, for which adding an up-down constraint improves peak detection accuracy, but makes the optimization problem more complicated.…
Efficiently reduces rank of non-negative matrices with quadratic time complexity.
We compare various extensions of the Bradley-Terry model and a hierarchical Poisson log-linear model in terms of their performance in predicting the outcome of soccer matches (win, draw, or loss). The parameters of the Bradley-Terry extensions are estimated by maximizing the log-likelihood, or an appropriately penalize…
McKernel introduces a framework to use kernel approximates in the mini-batch setting with Stochastic Gradient Descent (SGD) as an alternative to Deep Learning. Based on Random Kitchen Sinks [Rahimi and Recht 2007], we provide a C++ library for Large-scale Machine Learning. It contains a CPU optimized implementation of …
Log-linear models are a family of probability distributions which capture relationships between variables. They have been proven useful in a wide variety of fields such as epidemiology, economics and sociology. The interest in using these models is that they are able to capture context-specific independencies, relation…
We study surfaces in Euclidean space that are minimal for a log-linear density , where are real numbers not all zero. We prove that if a surface is -minimal foliated by circles in parallel planes, then these planes are orthogonal to the vector and the surface must…
A flexible nonparametric online changepoint detection algorithm for high-frequency data.
The paper integrates multiple Gaussian process predictions using Monte Carlo sampling.
Log-linear models are arguably the most successful class of graphical models for large-scale applications because of their simplicity and tractability. Learning and inference with these models require calculating the partition function, which is a major bottleneck and intractable for large state spaces. Importance Samp…
Neural models learn continuous-time Markov chain transition rates from data.
Paper explores duality in DPPs using embedding structure analysis.
In this paper, we classify the class of constant weighted curvature curves in the plane with a log-linear density, or in other words, classify all traveling curved fronts with a constant forcing term in The classification gives some interesting phenomena and consequences including: the family of curves conv…
HGConv uses HRR to efficiently detect malware, outperforming existing methods.
A major problem for the learning of Bayesian networks (BNs) is the exponential number of parameters needed for conditional probability tables. Recent research reduces this complexity by modeling local structure in the probability tables. We examine the use of log-linear local models. While log-linear models in this con…
New algorithm reconstructs sparse networks in subquadratic time.
Improved speech recognition with language model integration in sequence-to-sequence models.
Reducing ICD-10 code granularity improves cost model accuracy and stability.
In this paper we show that for the purposes of dimensionality reduction certain class of structured random matrices behave similarly to random Gaussian matrices. This class includes several matrices for which matrix-vector multiply can be computed in log-linear time, providing efficient dimensionality reduction of gene…
This paper is a continuation of Ishitani and Kato (2015), in which we derived a continuous-time value function corresponding to an optimal execution problem with uncertain market impact as the limit of a discrete-time value function. Here, we investigate some properties of the derived value function. In particular, we …
CANN models improve insurance claim count predictions using telematics data.
Hidden variables are ubiquitous in practical data analysis, and therefore modeling marginal densities and doing inference with the resulting models is an important problem in statistics, machine learning, and causal inference. Recently, a new type of graphical model, called the nested Markov model, was developed which …
We present a novel blind source separation (BSS) method, called information geometric blind source separation (IGBSS). Our formulation is based on the log-linear model equipped with a hierarchically structured sample space, which has theoretical guarantees to uniquely recover a set of source signals by minimizing the K…
New methods combine model predictions to avoid linear mixtures' limitations.
Learning a regression function using censored or interval-valued output data is an important problem in fields such as genomics and medicine. The goal is to learn a real-valued prediction function, and the training output labels indicate an interval of possible values. Whereas most existing algorithms for this task are…
Paper addresses online alignment of large language models under uncertain preference feedback.
A new probabilistic mixup framework improves deep learning generalization.
Privacy-preserving inference for clinical trials using differential privacy.
New method decomposes KL error using refined information and mode interactions.
Calculation of the log-normalizer is a major computational obstacle in applications of log-linear models with large output spaces. The problem of fast normalizer computation has therefore attracted significant attention in the theoretical and applied machine learning literature. In this paper, we analyze a recently pro…
A -translating soliton with density vector is a surface in Euclidean space whose mean curvature satisfies , where is the Gauss map. We classify all -translating solitons that are invariant by a one-parameter group of translations and a one-parameter group of rotat…
A new method for efficient BNC parameter estimation outperforms HDP smoothing.
We study the problem of collaborative filtering where ranking information is available. Focusing on the core of the collaborative ranking process, the user and their community, we propose new models for representation of the underlying permutations and prediction of ranks. The first approach is based on the assumption …
AFTNet uses a network-constrained Weibull model for biomarker discovery.
Estimates population size using capture-recapture designs with binary indicators.
We extend the theory of asymmetric information in mispricing models for stocks following geometric Brownian motion to constant relative risk averse investors. Mispricing follows a continuous mean--reverting Ornstein--Uhlenbeck process. Optimal portfolios and maximum expected log--linear utilities from terminal wealth f…
The thresholded feature has recently emerged as an extremely efficient, yet rough empirical approximation, of the time-consuming sparse coding inference process. Such an approximation has not yet been rigorously examined, and standard dictionaries often lead to non-optimal performance when used for computing thresholde…
Learning the Markov network structure from data is a problem that has received considerable attention in machine learning, and in many other application fields. This work focuses on a particular approach for this purpose called independence-based learning. Such approach guarantees the learning of the correct structure …
We introduce a new embarrassingly parallel parameter learning algorithm for Markov random fields with untied parameters which is efficient for a large class of practical models. Our algorithm parallelizes naturally over cliques and, for graphs of bounded degree, its complexity is linear in the number of cliques. Unlike…
This paper considers the problem of Byzantine fault tolerance in distributed linear regression in a multi-agent system. However, the proposed algorithms are given for a more general class of distributed optimization problems, of which distributed linear regression is a special case. The system comprises of a server and…
Scalable method for regionalizing and extracting temporal patterns from time series data.
A novel method for learning DAGs from positive-valued data.
We analyze a model of learning and belief formation in networks in which agents follow Bayes rule yet they do not recall their history of past observations and cannot reason about how other agents' beliefs are formed. They do so by making rational inferences about their observations which include a sequence of independ…