A new method detects and displays pairwise dependence between variates.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Many datasets are in the form of tables of binned data. Performing regression on these data usually involves either reading off bin heights, ignoring data from neighbouring bins or interpolating between bins thus over or underestimating the true bin integrals. In this paper we propose an elegant method for performing G…
Histogram binning method proven with guarantees without splitting data.
Isotonic regression binning affects calibration statistics of machine learning models.
Improved binning technique boosts nUV measure performance.
Balls-and-Bins sampling improves DP-SGD privacy and utility.
Bin Packing problems have been widely studied because of their broad applications in different domains. Known as a set of NP-hard problems, they have different vari- ations and many heuristics have been proposed for obtaining approximate solutions. Specifically, for the 1D variable sized bin packing problem, the two ke…
New methods reduce bias in estimating calibration error.
Improved kernel ridge regression for large datasets using weighted random binning.
The optimal binning is the optimal discretization of a variable into bins given a discrete or continuous numeric target. We present a rigorous and extensible mathematical programming formulation for solving the optimal binning problem for a binary, continuous and multi-class target type, incorporating constraints not p…
This paper improves multi-class calibration methods using mutual information maximization-based binning.
Paper analyzes ECE bias and provides bounds for its estimation.
New methods improve estimation of nonhomogeneous Poisson processes from limited data.
The MAP-Elites algorithm produces a set of high-performing solutions that vary according to features defined by the user. This technique has the potential to be a powerful tool for design space exploration, but is limited by the need for numerous evaluations. The Surrogate-Assisted Illumination algorithm (SAIL), introd…
New bin-wise scaling methods improve prediction uncertainty calibration for machine learning.
A method for non-parametric conditional distribution estimation using CRPS-optimal binning.
Solves online 3D bin packing with deep reinforcement learning under constraints.
Probability Density Estimation (PDE) is a multivariate discrimination technique based on sampling signal and background densities defined by event samples from data or Monte-Carlo (MC) simulations in a multi-dimensional phase space. In this paper, we present a modification of the PDE method that uses a self-adapting bi…
Study three types of uncertainty quantification for binary classification without distributional assumptions.
This paper introduces minimum-risk recalibration for probabilistic classifiers, improving their reliability and accuracy.
Object detection in streaming images is a major step in different detection-based applications, such as object tracking, action recognition, robot navigation, and visual surveillance applications. In mostcases, image quality is noisy and biased, and as a result, the data distributions are disturbed and imbalanced. Most…
A new survival analysis method eliminates hyperparameter tuning.
This study examines how discretization improves neural forecasting models.
New method unfolds distribution moments directly from data without binning.
In this paper we perform a statistical analysis over the returns and relative prices of the CAC and the S\&P with the purpose of analyzing the intra-day seasonalities of single and cross-sectional stock dynamics. In order to do that, we characterized the dynamics of a stock (or a set of stocks) by the evolut…
A new DP algorithm improves privacy in hashing and sampling for search and learning.
The objective of this work is to take advantage of deep neural networks in order to make next day crime count predictions in a fine-grain city partition. We make predictions using Chicago and Portland crime data, which is augmented with additional datasets covering weather, census data, and public transportation. The c…
For various applications, the relations between the dependent and independent variables are highly nonlinear. Consequently, for large scale complex problems, neural networks and regression trees are commonly preferred over linear models such as Lasso. This work proposes learning the feature nonlinearities by binning fe…
TCE measures calibration error with a test-based approach.
Recent developments have linked causal inference with Algorithmic Information Theory, and methods have been developed that utilize Conditional Kolmogorov Complexity to determine causation between two random variables. We present a method for inferring causal direction between continuous variables by using an MDL Binnin…
In subgroup discovery, also known as supervised pattern mining, discovering high quality one-dimensional subgroups and refinements of these is a crucial task. For nominal attributes, this is relatively straightforward, as we can consider individual attribute values as binary features. For numerical attributes, the task…
A 3D flexible bin packing problem (3D-FBPP) arises from the process of warehouse packing in e-commerce. An online customer's order usually contains several items and needs to be packed as a whole before shipping. In particular, 5% of tens of millions of packages are using plastic wrapping as outer packaging every day, …
Paper analyzes nonconvex bandit problems with improved adaptive methods.
Overconfidence and underconfidence in machine learning classifiers is measured by calibration: the degree to which the probabilities predicted for each class match the accuracy of the classifier on that prediction. How one measures calibration remains a challenge: expected calibration error, the most popular metric, ha…
Unified framework connects credit risk metrics with information theory.
Unified framework for PDF estimation using MDL-based binning and tensor factorization.
We consider the non-parametric regression problem under Huber's -contamination model, in which an fraction of observations are subject to arbitrary adversarial noise. We first show that a simple local binning median step can effectively remove the adversary noise and this median estimator is minimax optimal up t…
Applications such as weather forecasting and personalized medicine demand models that output calibrated probability estimates---those representative of the true likelihood of a prediction. Most models are not calibrated out of the box but are recalibrated by post-processing model outputs. We find in this work that popu…
HFNO enhances interpretability of turbulent flows through parallel wavenumber bin processing.
Paper predicts recycling bin full events to reduce RVM downtime.
In the framework of Multifractal Diffusion Entropy Analysis we propose a method for choosing an optimal bin-width in histograms generated from underlying probability distributions of interest. The method presented uses techniques of Rényi's entropy and the mean squared error analysis to discuss the conditions under whi…
AdaCat improves density estimation and planning in autoregressive models.
Bayesian inference engines improve density estimation accuracy and scalability.
New methods improve neural network extrapolation to long sequences.
Improved online algorithm for convex losses with near-optimal swap regret.
Positive-definite kernel functions are fundamental elements of kernel methods and Gaussian processes. A well-known construction of such functions comes from Bochner's characterization, which connects a positive-definite function with a probability distribution. Another construction, which appears to have attracted less…
A new method improves quantile regression for high-dimensional data.
Smart bin monitors predict medication adherence with high accuracy.