Proximal Mediation Analysis with Hidden Recanting Witnesses
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new witness two-sample test improves data efficiency and power.
AutoML simplifies two-sample tests for detecting distribution shifts.
We provide a new approach to training neural models to exhibit transparency in a well-defined, functional manner. Our approach naturally operates over structured data and tailors the predictor, functionally, towards a chosen family of (local) witnesses. The estimation problem is setup as a co-operative game between an …
The concepts of risk-aversion, chance-constrained optimization, and robust optimization have developed significantly over the last decade. Statistical learning community has also witnessed a rapid theoretical and applied growth by relying on these concepts. A modeling framework, called distributionally robust optimizat…
USD algorithm transports distributions with or without mass conservation.
In machine learning, we are given a dataset of the form , drawn as i.i.d. samples from an unknown probability distribution ; the marginal distribution for the 's being . We propose that rather than using a positive kernel such as the Gaussian for estimation of these…
Proposes GFMMD for comparing signals on graphs.
We present new excess risk bounds for general unbounded loss functions including log loss and squared loss, where the distribution of the losses may be heavy-tailed. The bounds hold for general estimators, but they are optimized when applied to -generalized Bayesian, MDL, and empirical risk minimization estimators. …
Computing Nash equilibrium (NE) of multi-player games has witnessed renewed interest due to recent advances in generative adversarial networks. However, computing equilibrium efficiently is challenging. To this end, we introduce the Gradient-based Nikaido-Isoda (GNI) function which serves: (i) as a merit function, vani…
OMLE combines optimism and MLE for efficient sequential decision making.
Geometric framework for signed multivariate tail-dependence compatibility at various thresholds.
Optimally estimates a functional using nuisance function tuning and sample splitting.
A new method for analyzing adaptive experiments using kernel treatment effects.
We prove optimal bounds for the convergence rate of ordinal embedding (also known as non-metric multidimensional scaling) in the 1-dimensional case. The examples witnessing optimality of our bounds arise from a result in additive number theory on sets of integers with no three-term arithmetic progressions. We also carr…
This note shows how to transform high-probability to in-expectation guarantees in machine learning.
New graphs show hierarchical hyperbolic properties, extending previous work.
A new method detects hidden driving forces in systems with multiple observables.
A new framework reduces RL sample complexity for complex MDPs.
We study the sample complexity of model-based reinforcement learning (henceforth RL) in general contextual decision processes that require strategic exploration to find a near-optimal policy. We design new algorithms for RL with a generic model class and analyze their statistical properties. Our algorithms have sample …
New methods estimate causal effects through mediators, handling confounding without strict assumptions.
Loxodromic elements are pseudo-Anosov on specific graphs.
Sharp comparison for sub-Gaussian random variables in convex order.
The past decade has witnessed a successful application of deep learning to solving many challenging problems in machine learning and artificial intelligence. However, the loss functions of deep neural networks (especially nonlinear networks) are still far from being well understood from a theoretical aspect. In this pa…
Face recall is a basic human cognitive process performed routinely, e.g., when meeting someone and determining if we have met that person before. Assisting a subject during face recall by suggesting candidate faces can be challenging. One of the reasons is that the search space - the face space - is quite large and lac…
Penalized estimation can conduct variable selection and parameter estimation simultaneously. The general framework is to minimize a loss function subject to a penalty designed to generate sparse variable selection. The majorization-minimization (MM) algorithm is a computational scheme for stability and simplicity, and …
We study a continuous-time version of the intermediation model of Grossman and Miller (1988). To wit, we solve for the competitive equilibrium prices at which liquidity takers' demands are absorbed by dealers with quadratic inventory costs, who can in turn gradually transfer these positions to an exogenous open market …
AUC (Area under the ROC curve) is an important performance measure for applications where the data is highly imbalanced. Learning to maximize AUC performance is thus an important research problem. Using a max-margin based surrogate loss function, AUC optimization problem can be approximated as a pairwise rankSVM learni…
In statistical learning theory, convex surrogates of the 0-1 loss are highly preferred because of the computational and theoretical virtues that convexity brings in. This is of more importance if we consider smooth surrogates as witnessed by the fact that the smoothness is further beneficial both computationally- by at…
Study compares various optimization algorithms for deep learning.
Link's sphere number equals its bridge number.
AdaDetectGPT improves text authorship detection with statistical guarantees.
Proposes DR-ME test for interpretable distributional treatment effects.
FedProx algorithm improved for non-smooth and heterogeneous data.
In this paper we show that certain generalizations of the -Whitney topology, which include the Hölder-Whitney and Sobolev-Whitney topologies on smooth manifolds, satisfy the Baire property, to wit, the countable intersection of open and dense sets is dense.
The study of networks has witnessed an explosive growth over the past decades with several ground-breaking methods introduced. A particularly interesting -- and prevalent in several fields of study -- problem is that of inferring a function defined over the nodes of a network. This work presents a versatile kernel-base…
We present a novel methodology based on a Taylor expansion of the network output for obtaining analytical expressions for the expected value of the network weights and output under stochastic training. Using these analytical expressions the effects of the hyperparameters and the noise variance of the optimization algor…
Recent years have witnessed a tremendous improvement of deep reinforcement learning. However, a challenging problem is that an agent may suffer from inefficient exploration, particularly for on-policy methods. Previous exploration methods either rely on complex structure to estimate the novelty of states, or incur sens…
Example shows dense subgroup of SL5(Z) not finitely presented.
Gaussian processes (GPs) have been proven to be powerful tools in various areas of machine learning. However, there are very few applications of GPs in the scenario of multi-view learning. In this paper, we present a new GP model for multi-view learning. Unlike existing methods, it combines multiple views by regularizi…
Nonparametric tests via kernel embedding of distributions have witnessed a great deal of practical successes in recent years. However, statistical properties of these tests are largely unknown beyond consistency against a fixed alternative. To fill in this void, we study here the asymptotic properties of goodness-of-fi…
Machine learning has witnessed tremendous success in solving tasks depending on a single hyperparameter. When considering simultaneously a finite number of tasks, multi-task learning enables one to account for the similarities of the tasks via appropriate regularizers. A step further consists of learning a continuum of…
NVGD uses neural networks to infer distributions without kernel choices.
Unified stability bounds for noisy SGD across convex and non-convex losses.
We consider the problem of configuring general-purpose solvers to run efficiently on problem instances drawn from an unknown distribution. The goal of the configurator is to find a configuration that runs fast on average on most instances, and do so with the least amount of total work. It can run a chosen solver on a r…
The last decade has witnessed an explosion in the development of models, theory and computational algorithms for "big data" analysis. In particular, distributed computing has served as a natural and dominating paradigm for statistical inference. However, the existing literature on parallel inference almost exclusively …
Survey on quantum computing and neural networks.
Unified approach for multicalibration in weakly supervised learning.