A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
In this paper, we consider a supervised learning setting where side knowledge is provided about the labels of unlabeled examples. The side knowledge has the effect of reducing the hypothesis space, leading to tighter generalization bounds, and thus possibly better generalization. We consider several types of side knowl…
Most existing distance metric learning methods assume perfect side information that is usually given in pairwise or triplet constraints. Instead, in many real-world applications, the constraints are derived from side information, such as users' implicit feedbacks and citations among articles. As a result, these constra…
The jet bundle description of time-dependent mechanics is revisited. The constraint algorithm for singular Lagrangians is discussed and an exhaustive description of the constraint functions is given. By means of auxiliary connections we give a basis of constraint functions in the Lagrangian and Hamiltonian sides. An ad…
In this paper, we propose a model-based clustering method (TVClust) that robustly incorporates noisy side information as soft-constraints and aims to seek a consensus between side information and the observed data. Our method is based on a nonparametric Bayesian hierarchical model that combines the probabilistic model …
In this paper, we propose a semi-supervised clustering method, CEC-IB, that models data with a set of Gaussian distributions and that retrieves clusters based on a partial labeling provided by the user (partition-level side information). By combining the ideas from cross-entropy clustering (CEC) with those from the inf…
Regulating causal effects through averaged constraints fails to enforce conditional independence.
problem Enforcing conditional independence in regulatory and analytic settings.
method Formulated causal masking as a linear program and analyzed the resulting enforcement problem from both regulator and optimizer perspectives.
result Averaged-constraint optimization often violates stratum-wise requirements while satisfying the averaged one exactly, and detection requires conditional-independence tests.
The one-bit quantization is implemented by one single comparator that operates at low power and a high rate. Hence one-bit compressive sensing (1bit-CS) becomes attractive in signal processing. When measurements are corrupted by noise during signal acquisition and transmission, 1bit-CS is usually modeled as minimizing …
This article provides a novel framework to evaluate limit order tactics that highlights expected fill price, adverse price selection cost, and opportunity cost. We formulate the problem of optimal execution of market orders with nonlinear market impact, power law decay kernel, and stochastic and deterministic liquidity…
We consider the problem of metric learning subject to a set of constraints on relative-distance comparisons between the data items. Such constraints are meant to reflect side-information that is not expressed directly in the feature vectors of the data items. The relative-distance constraints used in this work are part…
In the world of modern financial theory, portfolio construction has traditionally operated under at least one of two central assumptions: the constraints are derived from a utility function and/or the multivariate probability distribution of the underlying asset returns is fully known. In practice, both the performance…
We introduce a longevity feature to the classical optimal dividend problem by adding a constraint on the time of ruin of the firm. We extend the results in \cite{HJ15}, now in context of one-sided Lévy risk models. We consider de Finetti's problem in both scenarios with and without fix transaction costs, e.g. taxes. We…
Having a regression model, we are interested in finding two-sided intervals that are guaranteed to contain at least a desired proportion of the conditional distribution of the response variable given a specific combination of predictors. We name such intervals predictive intervals. This work presents a new method to fi…
Transposable data represents interactions among two sets of entities, and are typically represented as a matrix containing the known interaction values. Additional side information may consist of feature vectors specific to entities corresponding to the rows and/or columns of such a matrix. Further information may also…
ML Compass helps organizations choose AI models that balance utility, cost, and compliance.
problem Selecting AI models that meet user utility, deployment costs, and compliance requirements.
method Develops ML Compass, a framework for constrained optimization over a capability-cost frontier, using internal measures and empirical data.
result ML Compass produces deployment-aware recommendations that differ from capability-only rankings, clarifying trade-offs between capability, cost, and safety.
Most of metric learning approaches are dedicated to be applied on data described by feature vectors, with some notable exceptions such as times series, trees or graphs. The objective of this paper is to propose a metric learning algorithm that specifically considers relational data. The proposed approach can take benef…
We define a notion of Hempel distance for one-sided Heegaard splittings and show that the existence of alternate surfaces restricts distance for one-sided splittings in a manner similar to Hartshorn's and Scharlemann-Tomova's results for two-sided splittings. We also show that every geometrically compressible one-sided…