We reformulate Lehmer's question from 1933 and a question due to Schinzel and Zassenhaus from 1965 in terms of a comparison of the Mahler measures and the houses, respectively, of monic integer reciprocal and skew-reciprocal polynomials of the same degree. This entails that understanding the difference between orientat…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
System filters and ranks medical answers using pre-trained models.
End-to-end ASR error detection using audio-transcript entailment.
The paper proposes methods to control errors in language generation models using textual entailment.
Word embeddings provide point representations of words containing useful semantic information. We introduce multimodal word distributions formed from Gaussian mixtures, for multiple word meanings, entailment, and rich uncertainty information. To learn these distributions, we propose an energy-based max-margin objective…
New framework for consistent submodular maximization with insertions and deletions.
Approaches KL divergence for learning multi-sense word distributions.
Training modern deep learning models requires large amounts of computation, often provided by GPUs. Scaling computation from one GPU to many can enable much faster training and research progress but entails two complications. First, the training library must support inter-GPU communication. Depending on the particular …
Neurosymbolic predictors fail to model uncertainty under independence assumption.
Learning graph representations via low-dimensional embeddings that preserve relevant network properties is an important class of problems in machine learning. We here present a novel method to embed directed acyclic graphs. Following prior work, we first advocate for using hyperbolic spaces which provably model tree-li…
American options are financial instruments that can be exercised at any time before expiration. In this paper we study the problem of pricing this kind of derivatives within a framework in which some of the properties --volatility and dividend policy-- of the underlaying stock can change at a random instant of time, bu…
In this article we extend cutting and blowing up to the nonrational symplectic toric setting. This entails the possibility of cutting and blowing up for symplectic toric manifolds and orbifolds in nonrational directions.
By representing words with probability densities rather than point vectors, probabilistic word embeddings can capture rich and interpretable semantic information and uncertainty. The uncertainty information can be particularly meaningful in capturing entailment relationships -- whereby general words such as "entity" co…
Paper analyzes proper losses and their performance in machine learning tasks.
We study data-driven assistants that provide congestion forecasts to users of shared facilities (roads, cafeterias, etc.), to support coordination between them, and increase efficiency of such collective systems. Key questions are: (1) when and how much can (accurate) predictions help for coordination, and (2) which as…
By a theorem of Kirchhoff if the six sphere admits an almost complex structure then the seven sphere is parallelizable, more crucial, he exhibited an explicit global frame constructed out of the given almost complex structure. This result implicitly equips the seven sphere with a definite H-space multiplication. We pro…
Recurrent neural networks have become ubiquitous in computing representations of sequential data, especially textual data in natural language processing. In particular, Bidirectional LSTMs are at the heart of several neural models achieving state-of-the-art performance in a wide variety of tasks in NLP. However, BiLSTM…
The Trek Separation Theorem (Sullivant et al. 2010) states necessary and sufficient conditions for a linear directed acyclic graphical model to entail for all possible values of its linear coefficients that the rank of various sub-matrices of the covariance matrix is less than or equal to n, for any given n. In this pa…
The paper introduces Relative Bias to quantify LLM bias systematically.
Cost-effective method detects language model hallucinations.
R package for multi-objective model selection in statistics.
The last few years have seen a staggering number of empirical studies of the robustness of neural networks in a model of adversarial perturbations of their inputs. Most rely on an adversary which carries out local modifications within prescribed balls. None however has so far questioned the broader picture: how to fram…
New method uses LLMs to generate detailed scientific hypotheses.
This paper is an attempt at understanding the quantum-like dynamics of financial markets in terms of non-differentiable price-time continuum having fractal properties. The main steps of this development are the statistical scaling, the non-differentiability hypothesis, and the equations of motion entailed by this hypot…
We prove that a wide class of correlated stochastic volatility models exactly measure an empirical fact in which past returns are anticorrelated with future volatilities: the so-called ``leverage effect''. This quantitative measure allows us to fully estimate all parameters involved and it will entail a deeper study on…
A geometric analysis of protein folding, which complements many of the models in the literature, is presented. We examine the process from unfolded strand to the point where the strand becomes self-interacting. A central question is how it is possible that so many initial configurations proceed to fold to a unique fina…
Modern neural networks are often augmented with an attention mechanism, which tells the network where to focus within the input. We propose in this paper a new framework for sparse and structured attention, building upon a smoothed max operator. We show that the gradient of this operator defines a mapping from real val…
Paper proposes an ensemble approach to improve fairness in classifier decisions.
We study online prediction of bounded stationary ergodic processes. To do so, we consider the setting of prediction of individual sequences and build a deterministic regression tree that performs asymptotically as well as the best L-Lipschitz constant predictors. Then, we show why the obtained regret bound entails the …
The Killing operator on a Riemannian manifold is a linear differential operator on vector fields whose kernel provides the infinitesimal Riemannian symmetries. The Killing operator is best understood in terms of its prolongation, which entails some simple tensor identities. These simple identities can be viewed as aris…
Embedding methods which enforce a partial order or lattice structure over the concept space, such as Order Embeddings (OE) (Vendrov et al., 2016), are a natural way to model transitive relational data (e.g. entailment graphs). However, OE learns a deterministic knowledge base, limiting expressiveness of queries and the…
In this paper we prove explicit formulas for all Willmore surfaces of revolution and demonstrate their use in the discussion of the associated Dirichlet boundary value problems. It is shown by an explicit example that symmetric Dirichlet boundary conditions do in general not entail the symmetry of the surface. In addit…
Whatever information a deep neural network has gleaned from training data is encoded in its weights. How this information affects the response of the network to future data remains largely an open question. Indeed, even defining and measuring information entails some subtleties, since a trained network is a determinist…
We discuss the construction of Sp(2)Sp(1)-structures whose fundamental form is closed. In particular, we find 10 new examples of 8-dimensional nilmanifolds that admit an invariant closed 4-form with stabiliser Sp(2)Sp(1). Our constructions entail the notion of SO(4)-structures on 7-manifolds. We present a thorough inve…
Proposes a new neural head for asymmetric representation learning.
The mean field methods, which entail approximating intractable probability distributions variationally with distributions from a tractable family, enjoy high efficiency, guaranteed convergence, and provide lower bounds on the true likelihood. But due to requirement for model-specific derivation of the optimization equa…
Taylor expansions improve reinforcement learning policies.
This work explores the trade-offs between stability and accuracy in statistical estimation.
Model shows incentives in shared order book can lead to free-rider problem.
Paper characterizes causal graphs from hard interventions and proposes a learning algorithm.
Multi-hop inference is necessary for machine learning systems to successfully solve tasks such as Recognising Textual Entailment and Machine Reading. In this work, we demonstrate the effectiveness of adaptive computation for learning the number of inference steps required for examples of different complexity and that l…
Combines curvature descriptors with TDA for graph model evaluation.
Jacobi solved geodesics on triaxial ellipsoids.
Stable solutions to a specific equation are one-dimensional.
New proof of harmonic map uniqueness with analytic targets.
We show that the metrical connection can be introduced in the two-dimensional Finsler space such that entailed parallel transports along curves joining points of the underlying manifold keep the two-vector angle as well as the length of the tangent vector, thereby realizing isometries of tangent spaces under the parall…
Deep convolutional neural networks (CNNs) used in practice employ potentially hundreds of layers and ,s of nodes. Such network sizes entail significant computational complexity due to the large number of convolutions that need to be carried out; in addition, a large number of parameters needs to be learned and…
Jacobian regularization boosts neural network robustness without degrading generalization.