Improved kernel ridge regression for large datasets using weighted random binning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Positive-definite kernel functions are fundamental elements of kernel methods and Gaussian processes. A well-known construction of such functions comes from Bochner's characterization, which connects a positive-definite function with a probability distribution. Another construction, which appears to have attracted less…
Kernel method has been developed as one of the standard approaches for nonlinear learning, which however, does not scale to large data set due to its quadratic complexity in the number of samples. A number of kernel approximation methods have thus been proposed in the recent years, among which the random features metho…
Bin Packing problems have been widely studied because of their broad applications in different domains. Known as a set of NP-hard problems, they have different vari- ations and many heuristics have been proposed for obtaining approximate solutions. Specifically, for the 1D variable sized bin packing problem, the two ke…
Spectral clustering is one of the most effective clustering approaches that capture hidden cluster structures in the data. However, it does not scale well to large-scale problems due to its quadratic complexity in constructing similarity graphs and computing subsequent eigendecomposition. Although a number of methods h…
A new method detects and displays pairwise dependence between variates.
Estimates sample size for subgroup analysis in randomized experiments.
For various applications, the relations between the dependent and independent variables are highly nonlinear. Consequently, for large scale complex problems, neural networks and regression trees are commonly preferred over linear models such as Lasso. This work proposes learning the feature nonlinearities by binning fe…
The MAP-Elites algorithm produces a set of high-performing solutions that vary according to features defined by the user. This technique has the potential to be a powerful tool for design space exploration, but is limited by the need for numerous evaluations. The Surrogate-Assisted Illumination algorithm (SAIL), introd…
Paper introduces RPWithPrior for efficient label differential privacy in regression.
New bin-wise scaling methods improve prediction uncertainty calibration for machine learning.
In subgroup discovery, also known as supervised pattern mining, discovering high quality one-dimensional subgroups and refinements of these is a crucial task. For nominal attributes, this is relatively straightforward, as we can consider individual attribute values as binary features. For numerical attributes, the task…
Recent developments have linked causal inference with Algorithmic Information Theory, and methods have been developed that utilize Conditional Kolmogorov Complexity to determine causation between two random variables. We present a method for inferring causal direction between continuous variables by using an MDL Binnin…
Improved online algorithm for convex losses with near-optimal swap regret.
Study three types of uncertainty quantification for binary classification without distributional assumptions.
Many datasets are in the form of tables of binned data. Performing regression on these data usually involves either reading off bin heights, ignoring data from neighbouring bins or interpolating between bins thus over or underestimating the true bin integrals. In this paper we propose an elegant method for performing G…
Histogram binning method proven with guarantees without splitting data.
Cheap permutation tests speed up distribution testing without sacrificing accuracy.
Isotonic regression binning affects calibration statistics of machine learning models.
New method for privacy amplification without sampling for matrix factorization.
Improved binning technique boosts nUV measure performance.
A new DP algorithm improves privacy in hashing and sampling for search and learning.
Balls-and-Bins sampling improves DP-SGD privacy and utility.
OPORP combines permutation and random projection for efficient data vector compression.
New methods reduce bias in estimating calibration error.
Unified framework connects credit risk metrics with information theory.
The optimal binning is the optimal discretization of a variable into bins given a discrete or continuous numeric target. We present a rigorous and extensible mathematical programming formulation for solving the optimal binning problem for a binary, continuous and multi-class target type, incorporating constraints not p…
We propose a spatial diffuseness feature for deep neural network (DNN)-based automatic speech recognition to improve recognition accuracy in reverberant and noisy environments. The feature is computed in real-time from multiple microphone signals without requiring knowledge or estimation of the direction of arrival, an…
Smart bin monitors predict medication adherence with high accuracy.
This paper improves multi-class calibration methods using mutual information maximization-based binning.
Paper analyzes ECE bias and provides bounds for its estimation.
New methods improve estimation of nonhomogeneous Poisson processes from limited data.
The fast demographic growth, together with the concentration of the population in cities and the increasing amount of daily waste, are factors that push to the limit the ability of waste assimilation by Nature. Therefore, we need technological means to make an optimal management of the waste collection process, which r…
ABM automates feature engineering and variable selection for loss-based models.
Recent advances in statistical theory, together with advances in the computational power of computers, provide alternative methods to do mass-univariate hypothesis testing in which a large number of univariate tests, can be properly used to compare MEEG data at a large number of time-frequency points and scalp location…
A method for non-parametric conditional distribution estimation using CRPS-optimal binning.
Solves online 3D bin packing with deep reinforcement learning under constraints.
Probability Density Estimation (PDE) is a multivariate discrimination technique based on sampling signal and background densities defined by event samples from data or Monte-Carlo (MC) simulations in a multi-dimensional phase space. In this paper, we present a modification of the PDE method that uses a self-adapting bi…
This paper introduces minimum-risk recalibration for probabilistic classifiers, improving their reliability and accuracy.
Object detection in streaming images is a major step in different detection-based applications, such as object tracking, action recognition, robot navigation, and visual surveillance applications. In mostcases, image quality is noisy and biased, and as a result, the data distributions are disturbed and imbalanced. Most…
A new survival analysis method eliminates hyperparameter tuning.
This study examines how discretization improves neural forecasting models.
New method unfolds distribution moments directly from data without binning.
In this paper we perform a statistical analysis over the returns and relative prices of the CAC and the S\&P with the purpose of analyzing the intra-day seasonalities of single and cross-sectional stock dynamics. In order to do that, we characterized the dynamics of a stock (or a set of stocks) by the evolut…
Neural network based architectures used for sound recognition are usually adapted from other application domains, which may not harness sound related properties. The ConditionaL Neural Network (CLNN) is designed to consider the relational properties across frames in a temporal signal, and its extension the Masked Condi…
A new perfectly truthful calibration measure improves prediction reliability.
Paper uses conformal prediction for solar power forecasting in electricity markets.
The goal of lossy data compression is to reduce the storage cost of a data set while retaining as much information as possible about something () that you care about. For example, what aspects of an image contain the most information about whether it depicts a cat? Mathematically, this corresponds to finding…