A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Existing generalization theories analyze the generalization performance mainly based on the model complexity and training process. The ignorance of the task properties, which results from the widely used IID assumption, makes these theories fail to interpret many generalization phenomena or guide practical learning tas…
The problem of clustering is considered, for the case when each data point is a sample generated by a stationary ergodic process. We propose a very natural asymptotic notion of consistency, and show that simple consistent algorithms exist, under most general non-parametric assumptions. The notion of consistency is as f…
The problem of clustering is considered, for the case when each data point is a sample generated by a stationary ergodic process. We propose a very natural asymptotic notion of consistency, and show that simple consistent algorithms exist, under most general non-parametric assumptions. The notion of consistency is as f…
Directed graphical models provide a useful framework for modeling causal or directional relationships for multivariate data. Prior work has largely focused on identifiability and search algorithms for directed acyclic graphical (DAG) models. In many applications, feedback naturally arises and directed graphical models …
Let L be an exact Lagrangian submanifold inside the cotangent bundle of a closed manifold N. We prove that if N satisfies a mild homotopy assumption then the image of π_2(L) in π_2(N) has finite index. We make no assumption on the Maslov class of L, and we make no orientability assumptions. The homotopy assumption is e…
We study the problem of aggregation noisy labels. Usually, it is solved by proposing a stochastic model for the process of generating noisy labels and then estimating the model parameters using the observed noisy labels. A traditional assumption underlying previously introduced generative models is that each object has…
We obtain sharp quantitative Laplacian upper and lower estimates under no assumption on curvatures. As a result, we derive quantitative Laplacian, area and volume comparison theorems for tubes in Riemannian and Kähler manifolds under weak integral curvature assumptions. We also give some applications, such as a general…
This work addresses the following question: Under what assumptions on the data generating process can one infer the causal graph from the joint distribution? The approach taken by conditional independence-based causal discovery methods is based on two assumptions: the Markov condition and faithfulness. It has been show…
In this paper, we generalize the Cosmetic Surgery Conjecture to an n-cusped hyperbolic 3-manifold and prove it under the assumption of another well-known conjecture in number theory, so called the Zilber-Pink Conjecture. For n=1 and 2, we show them without the assumption.
Emputation learns imputation models guided by missingness assumptions.
problem Learning imputation models for missing data given observed data.
method Guided by specific missingness assumptions, Emputation trains a deep generative model to learn the extrapolation distribution of missing variables.
result The population minimizer of the emputation risk recovers the target extrapolation distribution under various identification assumptions.
There is a large body of work on convergence rates either in passive or active learning. Here we first outline some of the main results that have been obtained, more specifically in a nonparametric setting under assumptions about the smoothness of the regression function (or the boundary between classes) and the margin…
In the paper two important theorems about complete affine spheres are generalized to the case of statistical structures on abstract manifolds. The assumption about constant sectional curvature is replaced by the assumption that the curvature satisfies some inequalities.
Differentially Private algorithms often need to select the best amongst many candidate options. Classical works on this selection problem require that the candidates' goodness, measured as a real-valued score function, does not change by much when one person's data changes. In many applications such as hyperparameter o…
We consider the fundamental problem of learning a single neuron x↦σ(w⊤x) using standard gradient methods. As opposed to previous works, which considered specific (and not always realistic) input distributions and activation functions σ(⋅), we ask whether a more general result is attainable, under mi…
GALA framework learns invariant graph representations via environment augmentation with minimal assumptions.
problem Learning invariant graph representations from different environments without additional assumptions.
method Developed GALA framework with minimal assumptions of variation sufficiency and consistency. Uses an assistant model to differentiate graph environment changes.
result Extracting maximally invariant subgraphs to proxy predictions identifies underlying invariant subgraphs for successful out-of-distribution generalization.
There is a large body of work on convergence rates either in passive or active learning. Here we outline some of the results that have been obtained, more specifically in a nonparametric setting under assumptions about the smoothness and the margin noise. We also discuss the relative merits of these underlying assumpti…
The paper establishes general results in Lorentzian optimal transport theory.
problem Establishing strong duality and optimality conditions in Lorentzian optimal transport.
method Providing non-trivial assumptions on measures, characterizing optimality, and proving regularity results.
result Regularity results for c-convex functions and (weak) Kantorovich potentials do not extend to the Lorentzian setting, but under suitable assumptions, they are locally semconvex.
The Lookahead optimizer improves SGD's performance and generalization without restrictive assumptions.
problem Improving the generalization of SGD with Lookahead.
method A rigorous stability and generalization analysis of the Lookahead optimizer with minibatch SGD, leveraging on-average model stability.
result Derives generalization bounds for convex and strongly convex problems without the restrictive Lipschitzness assumption, demonstrating a linear speedup with batch size.
Generative source separation methods such as non-negative matrix factorization (NMF) or auto-encoders, rely on the assumption of an output probability density. Generative Adversarial Networks (GANs) can learn data distributions without needing a parametric assumption on the output density. We show on a speech source se…
We study generalizations of Reifenberg's Theorem for measures in Rn under assumptions on the Jones' β-numbers, which appropriately measure how close the support is to being contained in a subspace. Our main results, which holds for general measures without density assumptions, give effective measure bounds…
Learning a causal effect from observational data is not straightforward, as this is not possible without further assumptions. If hidden common causes between treatment X and outcome Y cannot be blocked by other measurements, one possibility is to use an instrumental variable. In principle, it is possible under some…
As an automatic method of determining model complexity using the training data alone, Bayesian linear regression provides us a principled way to select hyperparameters. But one often needs approximation inference if distribution assumption is beyond Gaussian distribution. In this paper, we propose a Bayesian linear reg…