Pareto's 80/20 rule follows a Gaussian distribution with twice the mean standard deviation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new learning rule consistently reduces error over data samples.
Paper extends transfer learning for decision rules, improving treatment rule estimation.
Calibrating a trading rule using a historical simulation (also called backtest) contributes to backtest overfitting, which in turn leads to underperformance. In this paper we propose a procedure for determining the optimal trading rule (OTR) without running alternative model configurations through a backtest engine. We…
Differentially private method for estimating individualized treatment rules.
This paper improves fraud prevention rule sets in fintech by generating diverse rules and finding Pareto-optimal subsets.
In this paper, we propose an adaptive stopping rule for kernel-based gradient descent (KGD) algorithms. We introduce the empirical effective dimension to quantify the increments of iterations in KGD and derive an implementable early stopping strategy. We analyze the performance of the adaptive stopping rule in the fram…
A method for collecting human supervision that combines rules and instance labels.
The paper argues that machine learning is a falsificationist process.
Adversarial training is a technique for training robust machine learning models. To encourage robustness, it iteratively computes adversarial examples for the model, and then re-trains on these examples via some update rule. This work analyzes the performance of adversarial training on linearly separable data, and prov…
In this paper, we study the problem of learning probabilistic logical rules for inductive and interpretable link prediction. Despite the importance of inductive link prediction, most previous works focused on transductive link prediction and cannot manage previously unseen entities. Moreover, they are black-box models …
This research adapts scoring rules for training survival models, improving predictive performance.
There has been significant recent work on the theory and application of randomized coordinate descent algorithms, beginning with the work of Nesterov [SIAM J. Optim., 22(2), 2012], who showed that a random-coordinate selection rule achieves the same convergence rate as the Gauss-Southwell selection rule. This result su…
We introduce a new weight-decay scaling rule to maintain sublayer gains across different widths in modern scale-invariant architectures.
Method integrates logical rules into neural multi-hop reasoning for drug repurposing.
Machine learning selects the best prediction rules from noisy data.
We present the design and implementation of a custom discrete optimization technique for building rule lists over a categorical feature space. Our algorithm produces rule lists with optimal training performance, according to the regularized empirical risk, with a certificate of optimality. By leveraging algorithmic bou…
Active inference framework improves -statistic estimation efficiency.
New learning rules for wide neural networks without backpropagation.
Generative models learn rules at different timescales, revealing a 'innovation window'.
The artificial neural network shows powerful ability of inference, but it is still criticized for lack of interpretability and prerequisite needs of big dataset. This paper proposes the Rule-embedded Neural Network (ReNN) to overcome the shortages. ReNN first makes local-based inferences to detect local patterns, and t…
The paper tackles causal rule discovery from observational data.
We propose three new robust aggregation rules for distributed synchronous Stochastic Gradient Descent~(SGD) under a general Byzantine failure model. The attackers can arbitrarily manipulate the data transferred between the servers and the workers in the parameter server~(PS) architecture. We prove the Byzantine resilie…
A new DP algorithm for weighted ERM protects sensitive data in predictive models.
Proposes methods to learn from biased samples, ensuring robust decision rules.
We propose a novel robust aggregation rule for distributed synchronous Stochastic Gradient Descent~(SGD) under a general Byzantine failure model. The attackers can arbitrarily manipulate the data transferred between the servers and the workers in the parameter server~(PS) architecture. We prove the Byzantine resilience…
This article introduces a framework to estimate the value of evidence-based decision making.
SIRUS creates interpretable rules from random forests for regression.
New learning rules achieve optimal sample complexity for weakly supervised classification.
We prove the statistical consistency of kernel Partial Least Squares Regression applied to a bounded regression learning problem on a reproducing kernel Hilbert space. Partial Least Squares stands out of well-known classical approaches as e.g. Ridge Regression or Principal Components Regression, as it is not defined as…
A nonparametric kernel-based method for realizing Bayes' rule is proposed, based on representations of probabilities in reproducing kernel Hilbert spaces. Probabilities are uniquely characterized by the mean of the canonical map to the RKHS. The prior and conditional probabilities are expressed in terms of RKHS functio…
This work proposes optimal decision rules for hierarchical classifiers to better align with evaluation metrics.
Interpretable classifiers have recently witnessed an increase in attention from the data mining community because they are inherently easier to understand and explain than their more complex counterparts. Examples of interpretable classification models include decision trees, rule sets, and rule lists. Learning such mo…
The problem of adaptive noisy clustering is investigated. Given a set of noisy observations , , the goal is to design clusters associated with the law of 's, with unknown density with respect to the Lebesgue measure. Since we observe a corrupted sample, a direct approach as the popular …
Time series forecasting models fail to consistently select the best model across different datasets.
Paper tackles unknown variances in best-arm identification.
Recently, several authors have advocated the use of rule learning algorithms to model multi-label data, as rules are interpretable and can be comprehended, analyzed, or qualitatively evaluated by domain experts. Many rule learning algorithms employ a heuristic-guided search for rules that model regularities contained i…
The 1/3 Financial Rule helps prevent household bankruptcy through balanced spending, savings, and debt repayment.
New findings show second-order scoring rules can't accurately represent epistemic uncertainty.
Efficient classifier error estimation without re-training.
Being able to model correlations between labels is considered crucial in multi-label classification. Rule-based models enable to expose such dependencies, e.g., implications, subsumptions, or exclusions, in an interpretable and human-comprehensible manner. Albeit the number of possible label combinations increases expo…
We empirically show the superiority of the equally weighted S\&P 500 portfolio over Sharpe's market capitalization weighted S\&P 500 portfolio. We proceed to consider the MaxMedian rule, a non-proprietary rule designed for the investor who wishes to do his/her own investing on a laptop with the purchase of only 20 stoc…
We propose an offline-online procedure for Fourier transform based option pricing. The method supports the acceleration of such essential tasks of mathematical finance as model calibration, real-time pricing, and, more generally, risk assessment and parameter risk estimation. We adapt the empirical magic point interpol…
We consider the setting of sequential prediction of arbitrary sequences based on specialized experts. We first provide a review of the relevant literature and present two theoretical contributions: a general analysis of the specialist aggregation rule of Freund et al. (1997) and an adaptation of fixed-share rules of He…
We consider the problem of structure learning for Gaifman models and learn relational features that can be used to derive feature representations from a knowledge base. These relational features are first-order rules that are then partially grounded and counted over local neighborhoods of a Gaifman model to obtain the …
Invariant Causal Set Covering Machines avoid spurious associations.
Learning to see through data is central to contemporary forms of algorithmic knowledge production. While often represented as a mechanical application of rules, making algorithms work with data requires a great deal of situated work. This paper examines how the often-divergent demands of mechanization and discretion ma…
Paper proposes a new method for SP with covariates using PADR and ERM.