MOCA uses modular attention to estimate causal effects from complex data.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We show how to restructure the counterparty risk faced by the originator of a securitization or covered bond arising from an interest rate hedging swap assisted by a "one-way" collateral agreement. This risk emerges when the swap is negotiated between the special purpose vehicle and a third party that covers itself thr…
Posterior sampling from diffusion models is computationally hard.
Estimates joint probability distribution from 1-way marginals using low-rank tensors and random projections.
In this paper it is proven that there is at most one way, up to isotopy, in which a connected, hyperbolic, orientable 3-manifold can fiber over the circle with monodromy in the Torelli group.
New findings show limitations in converting private learning to online learning efficiently.
Embeddings are ubiquitous in machine learning, appearing in recommender systems, NLP, and many other applications. Researchers and developers often need to explore the properties of a specific embedding, and one way to analyze embeddings is to visualize them. We present the Embedding Projector, a tool for interactive v…
The paper proposes methods to optimize pAUC for deep learning using DRO.
New method reduces computational cost for selective inference.
For any knot with genus one and unknotting number one, other than the figure-eight knot, we prove that there is exactly one way to unknot it by means of a crossing change. In the case of the figure-eight knot, we prove that there are precisely two unknotting crossing changes. The proof uses sutured manifold theory and …
Deep chest X-ray classifiers show bias in predicting diagnoses.
Classifies surfaces of section for Seifert fibrations.
Our interest in this paper is in the construction of symbolic explanations for predictions made by a deep neural network. We will focus attention on deep relational machines (DRMs, first proposed by H. Lodhi). A DRM is a deep network in which the input layer consists of Boolean-valued functions (features) that are defi…
I discuss certain applications of the Ricci flow in physics. I first review how it arises in the renormalization group (RG) flow of a nonlinear sigma model. I then review the concept of a Ricci soliton and recall how a soliton was used to discuss the RG flow of mass in 2-dimensions. I then present recent results obtain…
I show that solutions of the SU(infinity) Toda field equation generating a fixed Einstein-Weyl space are governed by a linear equation on the Einstein-Weyl space. From this, obstructions to the existence of Toda solutions generating a given Einstein-Weyl space are found. I also give a classification of Einstein-Weyl sp…
Adaptive Monte Carlo methods are recent variance reduction techniques. In this work, we propose a mathematical setting which greatly relaxes the assumptions needed by for the adaptive importance sampling techniques presented by Vazquez-Abad and Dufresne, Fu and Su, and Arouna. We establish the convergence and asymptoti…
A compact real analytic Riemannian manifold M admits a canonical complexification with plurisubharmonic exhaustion function satisfying the homogeneous complex Monge-Ampere equation, called a Grauert tube. From the point of view of complex analysis, several authors have considered whether a given complex manifold can ar…
The problem of biclustering consists of the simultaneous clustering of rows and columns of a matrix such that each of the submatrices induced by a pair of row and column clusters is as uniform as possible. In this paper we approximate the optimal biclustering by applying one-way clustering algorithms independently on t…
Let M be a simple 3-manifold with a toral boundary component partial_0 M. If Dehn filling M along partial_0 M one way produces a toroidal manifold and Dehn filling M along partial_0 M another way produces a boundary-reducible manifold, then we show that the absolute value of the intersection number on partial_0 M of th…
The paper simplifies symmetries in complex geometric structures.
Conditional independence testing is a key problem required by many machine learning and statistics tools. In particular, it is one way of evaluating the usefulness of some features on a supervised prediction problem. We propose a novel conditional independence test in a predictive setting, and show that it achieves bet…
In recent years several adversarial attacks and defenses have been proposed. Often seemingly robust models turn out to be non-robust when more sophisticated attacks are used. One way out of this dilemma are provable robustness guarantees. While provably robust models for specific -perturbation models have been dev…
We continue the study of statistical/computational tradeoffs in learning robust classifiers, following the recent work of Bubeck, Lee, Price and Razenshteyn who showed examples of classification tasks where (a) an efficient robust classifier exists, in the small-perturbation regime; (b) a non-robust classifier can be l…
We describe two techniques that significantly improve the running time of several standard machine-learning algorithms when data is sparse. The first technique is an algorithm that effeciently extracts one-way and two-way counts--either real or expected-- from discrete data. Extracting such counts is a fundamental step…
One way of producing explicit Riemannian 6-manifolds with holonomy SU(3) is by integrating a flow of SU(2)-structures on a 5-manifold, called the hypo evolution flow. In this paper we classify invariant hypo SU(2)-structures on nilpotent 5-dimensional Lie groups. We characterize the hypo evolution flow in terms of gaug…
One way to interpret smoothness of a measure in infinite dimensions is quasi-invariance of the measure under a class of transformations. Usually such settings lack a reference measure such as the Lebesgue or Haar measure, and therefore we can not use smoothness of a density with respect to such a measure. We describe h…
One way to obtain invariants of some Legendrian submanifolds in 1-jet spaces , equipped with the standard contact structure, is through the Morse theoretic technique of generating families. This paper extends the invariant of generating family cohomology by giving it a product . To define the product, moduli…
We consider the problem of learning classifiers for labeled data that has been distributed across several nodes. Our goal is to find a single classifier, with small approximation error, across all datasets while minimizing the communication between nodes. This setting models real-world communication bottlenecks in the …
Multi-output prediction deals with the prediction of several targets of possibly diverse types. One way to address this problem is the so called problem transformation method. This method is often used in multi-label learning, but can also be used for multi-output prediction due to its generality and simplicity. In thi…
Around 2008 N. Kawazumi and S. Zhang introduced a new fundamental numerical invariant for compact Riemann surfaces. One way of viewing the Kawazumi-Zhang invariant is as a quotient of two natural hermitian metrics with the same first Chern form on the line bundle of holomorphic differentials. In this paper we determine…
δ-CLUE generates diverse explanations for model uncertainty.
Deep networks have enabled reinforcement learning to scale to more complex and challenging domains, but these methods typically require large quantities of training data. An alternative is to use sample-efficient episodic control methods: neuro-inspired algorithms which use non-/semi-parametric models that predict valu…
Deep neural networks (DNNs) can easily fit a random labeling of the training data with zero training error. What is the difference between DNNs trained with random labels and the ones trained with true labels? Our paper answers this question with two contributions. First, we study the memorization properties of DNNs. O…
Continuous control tasks in reinforcement learning are important because they provide an important framework for learning in high-dimensional state spaces with deceptive rewards, where the agent can easily become trapped into suboptimal solutions. One way to avoid local optima is to use a population of agents to ensure…
New interpretation of attention in Transformers and Graph Attention Networks.
New approach improves multi-head attention by making heads less similar.
Regularizes attention scores in vision transformers using bootstrapping.
Random forests with attention and self-attention improve regression performance.
Aligns attention distributions for improved accuracy and robustness.
A framework for transformer attention layers derived from SVR.
Elliptical Attention improves transformer performance by focusing on contextually relevant features.
Kernel PCA explains self-attention mechanisms in deep learning models.
Study examines asset pricing using various attention models, finding global self-attention and sliding window sparse attention models perform well.
LARF improves random forests with attention mechanisms and contamination models.
One way to generalize the boundary Yamabe problem posed by Escobar is to ask if a given metric on a compact manifold with boundary can be conformally deformed to have vanishing -curvature in the interior and constant -curvature on the boundary. When restricting to the closure of the positive -cone, this is…
Gated attention improves performance by using a hierarchical mixture of experts.
Precision farming uses data analysis to optimize crop management.
Investigates the fundamental components of attention mechanisms.