The paper introduces a method for forecasting corporate sales growth using multiple reference variables.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Screening is the problem of finding a superset of the set of non-zero entries in an unknown p-dimensional vector β* given n noisy observations. Naturally, we want this superset to be as small as possible. We propose a novel framework for screening, which we refer to as Multiple Grouping (MuG), that groups variables, pe…
This paper solves the multiple reference model problem in RLHF with exact solutions and sample complexity guarantees.
New methods improve LLM preference optimization by intelligently weighting multiple reference models.
We investigate multiple testing and variable selection using the Least Angle Regression (LARS) algorithm in high dimensions under the assumption of Gaussian noise. LARS is known to produce a piecewise affine solution path with change points referred to as the knots of the LARS path. The key to our results is an express…
Enhances anomaly detection using multiple reference datasets.
New method identifies causal effects with categorical unobserved confounders.
Spectral variability is one of the major issue when conducting hyperspectral unmixing. Within a given image composed of some elementary materials (herein referred to as endmember classes), the spectral signature characterizing these classes may spatially vary due to intrinsic component fluctuations or external factors …
While generative adversarial networks (GAN) have been widely adopted in various topics, in this paper we generalize the standard GAN to a new perspective by treating realness as a random variable that can be estimated from multiple angles. In this generalized framework, referred to as RealnessGAN, the discriminator out…
A method estimates causal parameters using a latent variable recovery.
We present a novel method for hierarchical topic detection where topics are obtained by clustering documents in multiple ways. Specifically, we model document collections using a class of graphical models called hierarchical latent tree models (HLTMs). The variables at the bottom level of an HLTM are observed binary va…
In this work we examine some of the problems associated with the development of machine learning models with the objective to achieve robust generalization capabilities on common-task multiple-database scenarios. Referred to as the "database variability problem", we focus on a specific medical domain (sleep staging in …
We investigate a generalized stochastic model with the property known as mean reversion, that is, the tendency to relax towards a historical reference level. Besides this property, the dynamics is driven by multiplicative and additive Wiener processes. While the former is modulated by the internal behavior of the syste…
We formulate and solve a tensor model using a latent-variable approach.
This work extends entropic optimal transport to non-product reference couplings, focusing on Gaussian cases.
Neural population activity often exhibits rich variability and temporal structure. This variability is thought to arise from single-neuron stochasticity, neural dynamics on short time-scales, as well as from modulations of neural firing properties on long time-scales, often referred to as "non-stationarity". To better …
This article proposes a biologically inspired neurocomputational architecture which learns associations between words and referents in different contexts, considering evidence collected from the literature of Psycholinguistics and Neurolinguistics. The multi-layered architecture takes as input raw images of objects (re…
New method learns cell trajectories from multiple snapshots.
The learning of predictive models for data-driven decision support has been a prevalent topic in many fields. However, construction of models that would capture interactions among input variables is a challenging task. In this paper, we present a new preference learning approach for multiple criteria sorting with poten…
The paper analyzes optimal consumption with past spending maximum as a reference.
Study measures gender bias in machine translation using multiple reference points.
Proposes a method to combine datasets with missing values using Gaussian process latent variables.
Proposes a framework to fuse heterogeneous data sources for better modeling.
New method identifies causal variables from multi-node interventions, expanding on previous single-node approaches.
ReQuestNet simplifies 5G channel estimation with a unified model.
We consider the least-square linear regression problem with regularization by the l1-norm, a problem usually referred to as the Lasso. In this paper, we present a detailed asymptotic analysis of model consistency of the Lasso. For various decays of the regularization parameter, we compute asymptotic equivalents of the …
Proposes a new approach to MSDA by introducing latent covariate shift to handle varying label distributions.
A new convex model tackles noisy separable NMF with provable correctness.
We present a preference learning framework for multiple criteria sorting. We consider sorting procedures applying an additive value model with diverse types of marginal value functions (including linear, piecewise-linear, splined, and general monotone ones) under a unified analytical framework. Differently from the exi…
Extends geostatistical simulation method to handle multiple variables and large grids.
Maximum a posteriori (MAP) inference over discrete Markov random fields is a fundamental task spanning a wide spectrum of real-world applications, which is known to be NP-hard for general graphs. In this paper, we propose a novel semidefinite relaxation formulation (referred to as SDR) to estimate the MAP assignment. A…
In this paper, we propose a mixture of probabilistic partial canonical correlation analysis (MPPCCA) that extracts the Causal Patterns from two multivariate time series. Causal patterns refer to the signal patterns within interactions of two elements having multiple types of mutually causal relationships, rather than a…
We characterize and study variable importance (VIMP) and pairwise variable associations in binary regression trees. A key component involves the node mean squared error for a quantity we refer to as a maximal subtree. The theory naturally extends from single trees to ensembles of trees and applies to methods like rando…
New method identifies latent causal variables from observed data, overcoming indeterminacies.
We consider the least-square linear regression problem with regularization by the -norm, a problem usually referred to as the Lasso. In this paper, we first present a detailed asymptotic analysis of model consistency of the Lasso in low-dimensional settings. For various decays of the regularization parameter, w…
Study characterizes spike deconvolution basin for noisy data.
Paper generalizes tensor-train approximation for complex random variables.
Optimized biopharmaceutical seed train design reduces variability and saves time.
We propose a modification of linear discriminant analysis, referred to as compressive regularized discriminant analysis (CRDA), for analysis of high-dimensional datasets. CRDA is specially designed for feature elimination purpose and can be used as gene selection method in microarray studies. CRDA lends ideas from $\el…
Many widely studied graphical models with latent variables lead to nontrivial constraints on the distribution of the observed variables. Inspired by the Bell inequalities in quantum mechanics, we refer to any linear inequality whose violation rules out some latent variable model as a "hidden variable test" for that mod…
A new method treats all variables equally in fitting data.
Graphical models are commonly used to represent conditional dependence relationships between variables. There are multiple methods available for exploring them from high-dimensional data, but almost all of them rely on the assumption that the observations are independent and identically distributed. At the same time, o…
Researchers develop a method to infer reference measures from observed functionals.
This paper explores the following question: what kind of statistical guarantees can be given when doing variable selection in high-dimensional models? In particular, we look at the error rates and power of some multi-stage regression methods. In the first stage we fit a set of candidate models. In the second stage we s…
I studied what role the US stock markets and money markets have possibly played in the Gross Private Domestic Investment (GPDI) of the United States from the year 1959 to the year 2001, Gross Private Domestic Investment refers to the total amount of investment spending by businesses and firms located within the borders…
TSRGA scales multivariate linear regression for feature-distributed data.
The paper models musical motif transformations in Beethoven's works.
X-SHAP assesses multiplicative variable contributions in machine learning models.