Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

98197295393 · Jun 202019922001200920182026
48 results for feature binning

Improved kernel ridge regression for large datasets using weighted random binning.

problem Efficiently approximating kernel matrices for large-scale datasets.
method Introduced weighted random binning features for locality sensitive hashing.
result Weighted random binning features generate Gaussian processes of any desired smoothness.

This work proposes a method to learn nonlinear feature relations using non-convex regularized binned regression.

problem Learning feature nonlinearities in large scale complex problems.
method Binning feature values, finding the best fit in each quantile using non-convex regularized linear regression, enforcing smoothness via piecewise-constant/linear approximation, and selecting a sparse subset of features.
result The proposed algorithm achieves linear rate of convergence while requiring near-minimal number of samples, accurately learning feature nonlinearities.

Efficient random binning features improve kernel methods for large datasets.

problem Kernel methods' quadratic complexity limits their scalability to large datasets.
method Proposes and analyzes Random Binning (RB) features, showing faster convergence and parallelizability.
result RB features achieve faster convergence rates and parallelizability advantages compared to other random features.

FLEXI optimizes binning for better subgroup discovery in numerical and ordinal attributes.

problem Mining high quality subgroups from numerical attributes is challenging.
method FLEXI uses optimal binning to find high quality binary features for both numeric and ordinal attributes.
result FLEXI outperforms state of the art with up to 25 times improvement in subgroup quality.

A new method for scalable spectral clustering using random binning features.

problem Scalability issues in spectral clustering for large-scale problems.
method Random Binning features to accelerate similarity graph construction and eigendecomposition.
result Achieves similar accuracy to standard spectral clustering but with linear computational cost.

SAIL improves design exploration by reducing evaluations.

problem Limited evaluations in design space exploration.
method Integrates approximative models and intelligent sampling.
result Efficiently produces accurate models and diverse high-performing solutions.

New bin-wise scaling methods improve prediction uncertainty calibration for machine learning.

problem Improving prediction uncertainty calibration for machine learning regression.
method Adaptations of Binwise Variance Scaling (BVS) with alternative loss functions and feature-based binning.
result Improved adaptivity and consistency in prediction uncertainty calibration.

Paper explores Polya's characterization of positive-definite kernels and random feature maps.

problem Characterizing positive-definite kernels and their random feature maps.
method Study Polya's criterion and derive novel kernels; compare random Fourier and binning feature maps.
result Random binning feature map yields a closer Euclidean inner product to the kernel.

Study three types of uncertainty quantification for binary classification without distributional assumptions.

problem Uncertainty quantification for binary classification in a distribution-free setting.
method Established theorems connecting calibration, confidence intervals, and prediction sets for score-based classifiers.
result Distribution-free calibration is only possible using scoring functions that partition feature space into countably many sets.

Histogram binning method proven with guarantees without splitting data.

problem Proving theoretical guarantees for histogram binning without sample splitting.
method Using Markov property of order statistics to prove calibration guarantees for original method.
result Proves histogram binning has strong calibration guarantees without sample splitting.

Isotonic regression binning affects calibration statistics of machine learning models.

problem Isotonic regression binning introduces aleatoric uncertainty in calibration statistics.
method Calibration error statistics are recalibrated using isotonic regression, which produces stratified uncertainties.
result Stratified uncertainties lead to significant differences in bin-based calibration statistics.

Improved binning technique boosts nUV measure performance.

problem Improving the performance of the nUV measure in real applications.
method Introduced the nUV measure, provided theoretical optimal binning techniques, and proposed algorithms for approximate solutions.
result Approximate binning techniques show 4-13% increase in AUC scores with statistical significance.

Unified framework connects credit risk metrics with information theory.

problem Disconnection between industry-standard metrics and statistical theory.
method Unified information-theoretic framework, proving IV equals PSI, deriving standard errors, formalizing trade-off, automated binning with XGBoost.
result Unified framework connects IV and PSI, providing statistical foundation for metrics.

This paper improves multi-class calibration methods using mutual information maximization-based binning.

problem Calibration of deep neural network predictions, especially for small prior classes.
method I-Max concept for binning, shared class-wise calibration strategy.
result Improves multi-class ranking and calibration performance using a small calibration set.

New methods improve estimation of nonhomogeneous Poisson processes from limited data.

problem Estimating nonhomogeneous Poisson processes from limited data.
method Formulated as a learning generalization problem, proposed adaptive and data-driven binning methods.
result Improved estimation of nonhomogeneous Poisson processes with limited data.

ABM automates feature engineering and variable selection for loss-based models.

problem Improving model performance through better feature engineering and variable selection.
method ABM uses group and fused lasso regularization to automatically select cutting points and variables.
result ABM integrates feature engineering, variable selection, and model training.

Recent advances in statistical theory, together with advances in the computational power of computers, provide alternative methods to do mass-univariate hypothesis testing in which a large number of univariate tests, can be properly used to compare MEEG data at a large number of time-frequency points and scalp location…

2014-06-25abs ↗pdf ↗

A method for non-parametric conditional distribution estimation using CRPS-optimal binning.

problem Non-parametric conditional distribution estimation.
method Partitioning covariate-sorted observations into bins to minimize LOO-CRPS, selecting K by K-fold cross-validation of test CRPS.
result Produces narrower prediction intervals with near-nominal coverage compared to split-conformal competitors.

Paper tackles flexible bin packing for e-commerce, reducing costs.

problem Optimizing packing of cuboid items into bins with minimal surface area.
method Multi-task Selected Learning approach to generate item packing sequence and orientation.
result Selected Learning method achieves 5.47% cost reduction compared to greedy algorithms.

Solves online 3D bin packing with deep reinforcement learning under constraints.

problem Challenges of packing items immediately without information and constraints.
method Constrained deep reinforcement learning (DRL) with feasibility predictor.
result Significantly outperforms state-of-the-art methods in online 3D bin packing.

Method infers causal direction using data discretization and complexity calculation.

problem Determining causal direction between continuous variables.
method MDL Binning technique for data discretization and complexity calculation.
result Captures the shape of the data to determine causal direction.

This paper introduces minimum-risk recalibration for probabilistic classifiers, improving their reliability and accuracy.

problem Improving the reliability and accuracy of probabilistic classifiers.
method Minimum-risk recalibration within the MSE decomposition framework, analyzing UMB method and label shift adaptation.
result The optimal number of bins for UMB scales with n1/3n^{1/3}, resulting in a risk bound of approximately O(n2/3)O(n^{-2/3}).

Robust nonparametric regression method removes adversarial noise effectively.

problem Nonparametric regression under adversarial noise contamination.
method Local binning median followed by kernel smoothing or local polynomial regression.
result Minimax optimality over Hölder and Sobolev classes with arbitrary smoothness.

System separates sounds from mixtures without ground truth info.

problem Sound separation from multi-channel mixtures without labeled data.
method Deep clustering on multi-channel mixtures, projecting bins to spatially correlated clusters.
result Performance matches ground truth separation using only multi-channel mixtures.

This study examines how discretization improves neural forecasting models.

problem Improving predictive performance of neural forecasting models.
method Empirical investigation of data binning techniques on various neural forecasting architectures.
result Data binning almost always improves forecasting accuracy, but the type of binning is less important.

Estimates sample size for subgroup analysis in randomized experiments.

problem Determining sample size for accurate subgroup analysis.
method Turns inference problem into simultaneous inference, calculates sample size based on confidence level and margin of error.
result Allows inversion of sample size to feasible number of treatment arms or partition complexity.

Paper tackles noisy, biased object detection in cluttered environments.

problem Noisy and biased object detection in unknown cluttered environments.
method Incremental active semi-supervised learning (IASSL) combining batch-based active learning and bin-based semi-supervised learning.
result Superior performance compared to state-of-the-art object detection methods.

In this paper we perform a statistical analysis over the returns and relative prices of the CAC 4040 and the S\&P 500500 with the purpose of analyzing the intra-day seasonalities of single and cross-sectional stock dynamics. In order to do that, we characterized the dynamics of a stock (or a set of stocks) by the evolut…

2015-01-21abs ↗pdf ↗

A new DP algorithm improves privacy in hashing and sampling for search and learning.

problem Improving privacy in hashing and sampling for large-scale applications.
method Combines differential privacy with one permutation hashing and bin-wise consistent weighted sampling.
result Proposes DP-OPH and DP-BCWS algorithms that enhance privacy while maintaining utility.