Develops methods for causal inference in compositional data using instrumental variables.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper advances FL algorithms for composite optimization and statistical recovery.
New method for analyzing compositional data, addressing biases in summary statistics.
This paper introduces compositional data analysis for financial ratios, improving industry-level analysis.
This work is an analytical and numerical study of the composition of several fractals into one and of the relation between the composite dimension and the dimensions of the component fractals. In the case of composition of standard IFS with segments of equal size, the composite dimension can be expressed as a function …
Estimates kernel eigenvalues for compositional dot-product kernels.
Extends knockoff filter for composite null hypotheses in variable selection.
Paper analyzes stability and generalization of SCO algorithms.
New financial ratios using compositional data improve analysis of firm health.
Adapts Altman's model to compositional data for bankruptcy prediction.
AutoBayes simplifies variational inference by composing models and optimizing them.
In this paper we develop proximal methods for statistical learning. Proximal point algorithms are useful in statistics and machine learning for obtaining optimization solutions for composite functions. Our approach exploits closed-form solutions of proximal operators and envelope representations based on the Moreau, Fo…
Sublinearly structured DNNs achieve feature learning consistency for compositional functions.
Study introduces a variational approach for efficient KL divergence estimation in Dirichlet mixture models.
Many machine learning, statistical inference, and portfolio optimization problems require minimization of a composition of expected value functions (CEVF). Of particular interest is the finite-sum versions of such compositional optimization problems (FS-CEVF). Compositional stochastic variance reduced gradient (C-SVRG)…
A new geometry-preserving method for interpreting compositional data.
Under Markovian assumptions, we leverage a Central Limit Theorem (CLT) for the empirical measure in the test statistic of the composite hypothesis Hoeffding test so as to establish weak convergence results for the test statistic, and, thereby, derive a new estimator for the threshold needed by the test. We first show t…
Throughout music history, theorists have identified and documented interpretable rules that capture the decisions of composers. This paper asks, "Can a machine behave like a music theorist?" It presents MUS-ROVER, a self-learning system for automatically discovering rules from symbolic music. MUS-ROVER performs feature…
In statistical analysis, measuring a score of predictive performance is an important task. In many scientific fields, appropriate scores were tailored to tackle the problems at hand. A proper score is a popular tool to obtain statistically consistent forecasts. Furthermore, a mathematical characterization of the proper…
We formalize notions of robustness for composite estimators via the notion of a breakdown point. A composite estimator successively applies two (or more) estimators: on data decomposed into disjoint parts, it applies the first estimator on each part, then the second estimator on the outputs of the first estimator. And …
We propose a novel and flexible rank-breaking-then-composite-marginal-likelihood (RBCML) framework for learning random utility models (RUMs), which include the Plackett-Luce model. We characterize conditions for the objective function of RBCML to be strictly log-concave by proving that strict log-concavity is preserved…
Algorithmic fairness, and in particular the fairness of scoring and classification algorithms, has become a topic of increasing social concern and has recently witnessed an explosion of research in theoretical computer science, machine learning, statistics, the social sciences, and law. Much of the literature considers…
This paper addresses measurement errors in high-dimensional compositional data using a log-contrast model calibration approach.
In this paper the correlation between education, research and macroeconomic strength of countries at a global scale is analyzed on the basis of statistical data published by the UNIDO and OECD. It uses sets of composite indicators describing the economical performance and competitiveness as well as those relevant for h…
BOIS optimizes complex systems by leveraging structural knowledge.
We consider the composition optimization with two expected-value functions in the form of , { which formulates many important problems in statistical learning and machine learning such as solving Bellman equations in reinforcement l…
Develops a method for estimating networks and covariate associations in compositional data.
This paper studies robust estimation methods in high dimensions, comparing model-averaged and composite quantile estimators.
Fast detection of changepoints in linear regression models.
New statistical models for predicting ranked preferences from partial orders.
Unified algorithm for minimizing composite functions with flexible design.
We provide novel theoretical results regarding local optima of regularized -estimators, allowing for nonconvexity in both loss and penalty functions. Under restricted strong convexity on the loss and suitable regularity conditions on the penalty, we prove that \emph{any stationary point} of the composite objective f…
While statistics and machine learning offers numerous methods for ensuring generalization, these methods often fail in the presence of adaptivity---the common practice in which the choice of analysis depends on previous interactions with the same dataset. A recent line of work has introduced powerful, general purpose a…
Diffusion models learn hierarchical composition rules from data.
In many applications one may acquire a composition of several signals that may be corrupted by noise, and it is a challenging problem to reliably separate the components from one another without sacrificing significant details. Adding to the challenge, in a compressive sensing framework, one is given only an undersampl…
In this paper we provide a conceptual overview of latent variable models within a probabilistic modeling framework, an overview that emphasizes the compositional nature and the interconnectedness of the seemingly disparate models commonly encountered in statistical practice.
Paper tackles distributed linear regression with compositional covariates.
Generative model synthesizes complex data structures with composite and nested types.
This paper shows that scientific discovery can be efficiently learned via compositional function trees, reducing the sample complexity.
New method estimates effects of multiple nutrients on blood glucose.
In this paper, we propose a compositional nonparametric method in which a model is expressed as a labeled binary tree of nodes, where each node is either a summation, a multiplication, or the application of one of the basis functions to one of the covariates. We show that in order to recover a labeled bi…
We consider a group of mean-variance investors with mimicking desire such that each investor is willing to penalize deviations of his portfolio composition from compositions of other group members. Penalizing norm constraints are already applied for statistical improvement of Markowitz portfolio procedure in order to c…
We consider the stochastic composition optimization problem proposed in \cite{wang2017stochastic}, which has applications ranging from estimation to statistical and machine learning. We propose the first ADMM-based algorithm named com-SVR-ADMM, and show that com-SVR-ADMM converges linearly for strongly convex and Lipsc…
Three supervised learning methods for selecting logratios in compositional data analysis.
Simplifies efficient estimation via automatic differentiation and probabilistic programming.
New models for analyzing microbiome data with interactions.
Extracting automatically the complex set of features composing real high-dimensional data is crucial for achieving high performance in machine--learning tasks. Restricted Boltzmann Machines (RBM) are empirically known to be efficient for this purpose, and to be able to generate distributed and graded representations of…
New algorithms solve nonconvex federated learning problems efficiently.