Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

2605217811,041 · Jun 202019922001200920182026
48 results for data-driven clause generation

Two scalable methods for PSL structure learning improve runtime and AUC.

problem Efficiently learning clauses for probabilistic soft logic models.
method Greedy search and a novel optimization method combining data-driven clause generation and PPLL objective.
result PPLL achieves up to 15% AUC gains and an order of magnitude runtime speedup.

ClauseLens uses reinforcement learning to price reinsurance treaties transparently and auditably.

problem Opaque and difficult-to-audit reinsurance treaty pricing practices.
method ClauseLens models treaty pricing as a Risk-Aware Constrained Markov Decision Process (RA-CMDP), incorporating legal clauses and generating interpretable explanations.
result ClauseLens reduces solvency violations and improves tail-risk performance, achieving 88.2% accuracy in clause-grounded explanations.

Faster Tsetlin Machines use clause indexing to speed inference and learning.

problem Overfitting and slow inference in Tsetlin Machines.
method Introduced a look-up table that indexes clauses based on feature falsification, enabling faster evaluation of clauses.
result Up to 15 times faster classification and three times faster learning on MNIST and Fashion-MNIST.

New study shows low-degree polynomial algorithms struggle at clause densities close to Fix's.

problem Finding satisfying assignments in random k-SAT formulas at high clause densities.
method Analysis of low-degree polynomial algorithms and a new many-way overlap gap property.
result No efficient algorithms can find satisfying assignments at clause densities close to Fix's.

CTM uses conjunctive clauses for image recognition, achieving high accuracy.

problem High computational complexity and lack of interpretability in CNNs.
method Introduces Convolutional Tsetlin Machine (CTM) using conjunctive clauses in propositional logic.
result CTM achieves competitive accuracy on various benchmarks, including MNIST and Fashion-MNIST.

Improved RTM uses integer weights to reduce computation and increase interpretability.

problem Lack of interpretability in nonlinear regression models.
method Integer weighted RTM clauses, combined with a novel learning scheme.
result Significantly reduced computation cost with improved accuracy.

This paper works out fair values of stock loan model with automatic termination clause, cap and margin. This stock loan is treated as a generalized perpetual American option with possibly negative interest rate and some constraints. Since it helps a bank to control the risk, the banks charge less service fees compared …

2010-05-09abs ↗pdf ↗

The paper introduces closed-form expressions for interpreting Tsetlin Machines.

problem Interpreting complex Tsetlin Machines with a large number of clauses.
method Developed closed-form expressions for local and global interpretability of Tsetlin Machines.
result The expressions enable real-time feature importance assessment and data clustering.

Pricing Chinese convertible bonds using Monte Carlo simulation and dynamic programming.

problem Pricing Chinese convertible bonds accurately.
method Monte Carlo simulation and dynamic programming with regression and backward induction.
result An underpriced strategy significantly outperforms benchmarks.

We present a case-study demonstrating the usefulness of Bayesian hierarchical mixture modelling for investigating cognitive processes. In sentence comprehension, it is widely assumed that the distance between linguistic co-dependents affects the latency of dependency resolution: the longer the distance, the longer the …

2017-02-02abs ↗pdf ↗

Insurance contracts for autonomous AI agents must be actuarially sound and resistant to gaming.

problem Designing insurance contracts for autonomous AI agents that are actuarially sound and resistant to gaming.
method Characterizing a five-attack space and proving the actuarial runtime is gaming-resistant.
result An incentive-compatible layer for actuarial control of autonomous-agent side effects.

In this note we show how to replicate a stylized CDS with a repurchase agreement and an asset swap. The latter must be designed in such a way that, on default of the issuer, it is terminated with a zero close-out amount. This break clause can be priced using the well known unilateral credit/debit valuation adjustment f…

2013-04-30abs ↗pdf ↗

Seglearn is an open-source python package for machine learning time series or sequences using a sliding window segmentation approach. The implementation provides a flexible pipeline for tackling classification, regression, and forecasting problems with multivariate sequence and contextual data. This package is compatib…

2018-03-21abs ↗pdf ↗

Paper characterizes and represents pairwise causal background knowledge for improved causal inference.

problem Improving causal inference by handling pairwise causal constraints.
method Graphical characterization, direct causal clause (DCC), unified representation, MPDAG, polynomial-time algorithms.
result Pairwise causal background knowledge uniquely decomposes into MPDAG and DCCs, improving causal effect identification.

Normal surface theory, a tool to represent surfaces in a triangulated 3-manifold combinatorially, is ubiquitous in computational 3-manifold theory. In this paper, we investigate a relaxed notion of normal surfaces where we remove the quadrilateral conditions. This yields normal surfaces that are no longer embedded. We …

2014-12-16abs ↗pdf ↗

New framework for data-driven hyperparameter tuning with structured loss.

problem Statistical foundations for multi-dimensional hyperparameter tuning remain limited.
method General framework using real algebraic geometry for semi-algebraic function classes.
result First general guarantees for multi-dimensional hyperparameter tuning.

The paper analyzes how CNNs interpret NLP tasks and identify linguistic features.

problem Understanding how CNNs capture linguistic features in NLP tasks.
method Visualization techniques and error analysis to interpret CNNs.
result Identified how CNNs capture different linguistic features and their impact on model performance.

Enhances data-driven models with physics knowledge for better system dynamics.

problem Improving generalization and interpretability in complex physical system modeling.
method EVGP (Explicit Variational Gaussian Process) model that incorporates domain knowledge into data-driven models.
result The EVGP model outperforms purely data-driven models when using prior domain knowledge.

A Python tool assesses fairness, accountability, and transparency in AI decisions.

problem Lack of regulation and certification for AI-driven decisions.
method Developed an open-source Python toolbox to analyze fairness, accountability, and transparency aspects of machine learning.
result Automatically reports fairness, accountability, and transparency aspects of AI decisions to stakeholders.

This paper bounds errors in data-driven power grid models using Rademacher complexity.

problem Ensuring accuracy of data-driven power grid models under incomplete physical information.
method Rademacher complexity theory for error bounds and evaluation implementation.
result Generalization error bounds for branch flow linearization and external network equivalent models.

Hybridizes physical and data-driven methods for predicting physicochemical properties.

problem Predicting physicochemical properties accurately using limited data.
method Distills physical method predictions into a prior model and combines with sparse experimental data using Bayesian inference.
result Significant improvements in predicting activity coefficients at infinite dilution compared to baselines and ensemble methods.

Study integrates machine learning with SAA for optimizing decisions based on uncertain parameters and covariates.

problem Optimizing decisions under uncertain parameters and covariates.
method Data-driven frameworks integrating machine learning prediction models within SAA for scenario generation.
result Consistent and asymptotically optimal solutions under certain conditions, with finite sample guarantees.

Neurally-Guided Structure Inference combines search and data-driven methods for efficient, robust structure inference.

problem Combining the advantages of exhaustive search and data-driven methods for structure inference.
method Neurally-Guided Structure Inference (NG-SI) uses a neural network to guide hierarchical search over structures.
result NG-SI outperforms search-based and data-driven methods on probabilistic matrix decomposition and symbolic program parsing.

A new method for support vector regression using a data-driven insensitive parameter.

problem Determining an optimal insensitive parameter in support vector regression.
method A data-driven approach to approximate the insensitive parameter by minimizing a generalized loss function based on the likelihood principle.
result The proposed method outperforms traditional support vector regression methods and has lower computational costs.

Lifted Relational Neural Networks (LRNNs) describe relational domains using weighted first-order rules which act as templates for constructing feed-forward neural networks. While previous work has shown that using LRNNs can lead to state-of-the-art results in various ILP tasks, these results depended on hand-crafted ru…

2017-10-05abs ↗pdf ↗

The MEM method uses data-driven priors for linear inverse problems, proving convergence and estimating differences.

problem Linear inverse problems with approximate priors.
method Maximum Entropy on the Mean (MEM) method with data-driven priors.
result Empirical mean convergence and estimates for prior differences based on epigraphical distance.

Algorithm extracts non-monotonic rules from statistical models using HUIM.

problem Extracting non-monotonic rules from statistical learning models.
method Reduces problem to HUIM, uses TreeExplainer for feature importance.
result Significant improvement in classification metrics and training time.

InVAErt networks use data-driven methods for system synthesis and identifiability analysis.

problem Model synthesis and identifiability analysis for complex systems.
method Deterministic encoder and decoder, normalizing flow, variational encoder, loss function penalty coefficients, latent space sampling.
result Validation through various system types, demonstrating effectiveness of the framework.

RCUKF combines data-driven modeling and Bayesian estimation for accurate system state estimation.

problem Challenges in obtaining reliable process models for complex systems.
method Integrates reservoir computing with unscented Kalman filtering.
result Demonstrated effectiveness on benchmark problems and real-time vehicle trajectory estimation.

Data-driven Distributionally Robust Optimization (DD-DRO) via optimal transport has been shown to encompass a wide range of popular machine learning algorithms. The distributional uncertainty size is often shown to correspond to the regularization parameter. The type of regularization (e.g. the norm used to regularize)…

2017-05-19abs ↗pdf ↗

GNPs learn operators on non-Euclidean geometries using neural networks.

problem Learning operators on complex geometries like manifolds.
method Geometric Neural Operators (GNPs) that incorporate geometric properties.
result GNPs can estimate metrics, solve PDEs, and learn LB operators on manifolds.

Data-driven method for error estimation without needing class complexity.

problem Constructing confidence intervals for a class of estimates.
method Data-driven approach to derive high-probability upper bounds on maximum error.
result Method naturally adapts to unknown correlation structures and works for finite and infinite classes.