We show that derivations of the differential structure of a subcartesian space satisfy the chain rule and have maximal integral curves.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper explores subdifferential chain rules for matrix factorization and related machine learning models.
In many healthcare settings, intuitive decision rules for risk stratification can help effective hospital resource allocation. This paper introduces a novel variant of decision tree algorithms that produces a chain of decisions, not a general tree. Our algorithm, -Carving Decision Chain (ACDC), sequentially carves o…
Although for neural networks with locally Lipschitz continuous activation functions the classical derivative exists almost everywhere, the standard chain rule is in general not applicable. We will consider a way of introducing a derivative for neural networks that admits a chain rule, which is both rigorous and easy to…
We derive generalization and excess risk bounds for neural nets using a family of complexity measures based on a multilevel relative entropy. The bounds are obtained by introducing the notion of generated hierarchical coverings of neural nets and by using the technique of chaining mutual information introduced in Asadi…
This paper presents a new methodology to compute first-order Greeks for barrier options under the framework of path-dependent payoff functions with European, Lookback, or Asian type and with time-dependent trigger levels. In particular, we develop chain rules for Wiener path integrals between two curves that arise in t…
CT compares two distributions using Bayes' theorem and chain rule.
Method measures weight similarity in neural networks using normalization and statistical inference.
It is proved that the members of the Riccati hierarchy, the so-called Riccati chain equations, can be considered as particular cases of projective Riccati equations, which greatly simplifies the study of the Riccati hierarchy. This also allows us to characterize Riccati chain equations geometrically in terms of the pro…
The Widrow-Hoff rule simplifies language data simulation.
Random Intersection Chains selects important interactions from categorical features.
A conformal procedure improves CoT reasoning by aggregating reasoning paths and calibrating abstention rules.
Paper finds optimal selling rule for pairs trading with stock constraints.
A new stopping rule based on E-values helps efficiently use sampling in Bayesian Deep Ensembles.
This paper is concerned with an optimal stock selling rule under a Markov chain model. The objective is to find an optimal stopping time to sell the stock so as to maximize an expected return. Solutions to the associated variational inequalities are obtained. Closed-form solutions are given in terms of a set of thresho…
Training activation quantized neural networks involves minimizing a piecewise constant function whose gradient vanishes almost everywhere, which is undesirable for the standard back-propagation or chain rule. An empirical way around this issue is to use a straight-through estimator (STE) (Bengio et al., 2013) in the ba…
We aim to predict and explain service failures in supply-chain networks, more precisely among last-mile pickup and delivery services to customers. We analyze a dataset of 500,000 services using (1) supervised classification with Random Forests, and (2) Association Rules. Our classifier reaches an average sensitivity of…
Backpropagation and the chain rule of derivatives have been prominent; however, the total derivative rule has not enjoyed the same amount of attention. In this work we show how the total derivative rule leads to an intuitive visual framework for creating gradient estimators on graphical models. In particular, previous …
A one-to-one correspondence is drawn between law invariant risk measures and divergences, which we define as functionals of pairs of probability measures on arbitrary standard Borel spaces satisfying a few natural properties. Divergences include many classical information divergence measures, such as relative entropy a…
We propose a new statistical model for computational linguistics. Rather than trying to estimate directly the probability distribution of a random sentence of the language, we define a Markov chain on finite sets of sentences with many finite recurrent communicating classes and define our language model as the invarian…
We study an extension of the classic stochastic multi-armed bandit problem which involves multiple plays and Markovian rewards in the rested bandits setting. In order to tackle this problem we consider an adaptive allocation rule which at each stage combines the information from the sample means of all the arms, with t…
In this paper we propose a novel approach for learning from data using rule based fuzzy inference systems where the model parameters are estimated using Bayesian inference and Markov Chain Monte Carlo (MCMC) techniques. We show the applicability of the method for regression and classification tasks using synthetic data…
New framework improves stochastic optimization for variational inference.
The paper develops a stationary-distribution theory for Random Forest ensemble size selection.
This paper analyzes voter coalitions in MakerDAO's decentralized governance.
Class imbalance is an intrinsic characteristic of multi-label data. Most of the labels in multi-label data sets are associated with a small number of training examples, much smaller compared to the size of the data set. Class imbalance poses a key challenge that plagues most multi-label learning methods. Ensemble of Cl…
New Sasaki-Einstein 7-spheres found via Berglund-Hübsch transpose.
New method simplifies causal inference with tiered background knowledge.
Optimal sample complexity for autoregressive chain-of-thought learning proven.
Proteins are linear molecular chains that often fold to function. The topology of folding is widely believed to define its properties and function, and knot theory has been applied to study protein structure and its implications. More that 97% of proteins are, however, classified as unknots when intra-chain interaction…
Monte Carlo methods are essential tools for Bayesian inference. Gibbs sampling is a well-known Markov chain Monte Carlo (MCMC) algorithm, extensively used in signal processing, machine learning, and statistics, employed to draw samples from complicated high-dimensional posterior distributions. The key point for the suc…
Improves sampling quality in model composition using MH-like acceptance rule for score-based diffusion models.
Upper bound on expected supremum of Bernoulli process.
New calculus on spacetimes for nonlinear differential equations.
Method learns CTMC models from steady-state data, predicting unseen states.
Two Heegaard Floer knot complexes are called stably equivalent if an acyclic complex can be added to each complex to make them filtered chain homotopy equivalent. Hom showed that if two knots are concordant, then their knot complexes are stably equivalent. Invariants of stable equivalence include the concordance invari…
LLMs translate natural language trading intents into correct option strategies using a domain-specific language.
ToolChain-CRC addresses the risk-control problem for retrieval-augmented and tool-using agents under drift.
Paper bridges matching rules and height functions in aperiodic tilings.
ISOMORPH creates a digital twin for supply chain logistics, advancing time-series forecasting benchmarks.
We define notions of differentiability for maps from and to the space of persistence barcodes. Inspired by the theory of diffeological spaces, the proposed framework uses lifts to the space of ordered barcodes, from which derivatives can be computed. The two derived notions of differentiability (respectively from and t…
This paper constructs an algebra on a 3-torus with specific properties for fluid dynamics.
Dynamic abstention improves LLM accuracy by selectively terminating unpromising reasoning.
New GAN formulation addresses mode collapse issue.
Improved Bayesian inference for neuronal ensemble inference reduces computational cost.
We propose a new method of discovering causal relationships in temporal data based on the notion of causal compression. To this end, we adopt the Pearlian graph setting and the directed information as an information theoretic tool for quantifying causality. We introduce chain rule for directed information and use it to…
We define an algebraic/combinatorial object on the front projection of a Legendrian knot called a Morse complex sequence, abbreviated MCS. This object is motivated by the theory of generating families and provides new connections between generating families, normal rulings, and augmentations of the Chekanov-Eliashb…
IterefinE combines KG refinement with embeddings to improve KG quality.