Improved kernel ridge regression for large datasets using weighted random binning.
problem Efficiently approximating kernel matrices for large-scale datasets.
method Introduced weighted random binning features for locality sensitive hashing.
result Weighted random binning features generate Gaussian processes of any desired smoothness.
Deep learning optimizes variable sized bin packing problems.
problem Optimizing bin packing for variable sized items.
method Deep learning approach with auto feature selection and accurate labeling.
result Deep learning system achieves high accuracy in selecting optimal heuristics.
This work proposes a method to learn nonlinear feature relations using non-convex regularized binned regression.
problem Learning feature nonlinearities in large scale complex problems.
method Binning feature values, finding the best fit in each quantile using non-convex regularized linear regression, enforcing smoothness via piecewise-constant/linear approximation, and selecting a sparse subset of features.
result The proposed algorithm achieves linear rate of convergence while requiring near-minimal number of samples, accurately learning feature nonlinearities.
Efficient random binning features improve kernel methods for large datasets.
problem Kernel methods' quadratic complexity limits their scalability to large datasets.
method Proposes and analyzes Random Binning (RB) features, showing faster convergence and parallelizability.
result RB features achieve faster convergence rates and parallelizability advantages compared to other random features.
FLEXI optimizes binning for better subgroup discovery in numerical and ordinal attributes.
problem Mining high quality subgroups from numerical attributes is challenging.
method FLEXI uses optimal binning to find high quality binary features for both numeric and ordinal attributes.
result FLEXI outperforms state of the art with up to 25 times improvement in subgroup quality.
A new method for scalable spectral clustering using random binning features.
problem Scalability issues in spectral clustering for large-scale problems.
method Random Binning features to accelerate similarity graph construction and eigendecomposition.
result Achieves similar accuracy to standard spectral clustering but with linear computational cost.
SAIL improves design exploration by reducing evaluations.
problem Limited evaluations in design space exploration.
method Integrates approximative models and intelligent sampling.
result Efficiently produces accurate models and diverse high-performing solutions.
New bin-wise scaling methods improve prediction uncertainty calibration for machine learning.
problem Improving prediction uncertainty calibration for machine learning regression.
method Adaptations of Binwise Variance Scaling (BVS) with alternative loss functions and feature-based binning.
result Improved adaptivity and consistency in prediction uncertainty calibration.
Paper explores Polya's characterization of positive-definite kernels and random feature maps.
problem Characterizing positive-definite kernels and their random feature maps.
method Study Polya's criterion and derive novel kernels; compare random Fourier and binning feature maps.
result Random binning feature map yields a closer Euclidean inner product to the kernel.
Study three types of uncertainty quantification for binary classification without distributional assumptions.
problem Uncertainty quantification for binary classification in a distribution-free setting.
method Established theorems connecting calibration, confidence intervals, and prediction sets for score-based classifiers.
result Distribution-free calibration is only possible using scoring functions that partition feature space into countably many sets.
Proposes a method for Gaussian Process regression on binned data.
problem Regression on binned data often leads to inaccurate predictions.
method Gaussian Process regression tailored for binned data.
result Effective probabilistic predictions of latent functions.
Histogram binning method proven with guarantees without splitting data.
problem Proving theoretical guarantees for histogram binning without sample splitting.
method Using Markov property of order statistics to prove calibration guarantees for original method.
result Proves histogram binning has strong calibration guarantees without sample splitting.
Isotonic regression binning affects calibration statistics of machine learning models.
problem Isotonic regression binning introduces aleatoric uncertainty in calibration statistics.
method Calibration error statistics are recalibrated using isotonic regression, which produces stratified uncertainties.
result Stratified uncertainties lead to significant differences in bin-based calibration statistics.
Improved binning technique boosts nUV measure performance.
problem Improving the performance of the nUV measure in real applications.
method Introduced the nUV measure, provided theoretical optimal binning techniques, and proposed algorithms for approximate solutions.
result Approximate binning techniques show 4-13% increase in AUC scores with statistical significance.
Optimal binning method for numeric targets using mathematical programming.
problem Optimizing the discretization of numeric variables for classification.
method Mathematical programming formulation for binary, continuous, and multi-class targets with constraints.
result Convex mixed-integer programming formulations for all target types.
Balls-and-Bins sampling improves DP-SGD privacy and utility.
problem Improving privacy and utility in DP-SGD implementations.
method Introducing Balls-and-Bins sampling as an alternative to shuffling in DP-SGD.
result Balls-and-Bins sampling achieves utility comparable to shuffling while offering better privacy amplification.
BIN-CT optimizes waste collection routes to reduce costs and emissions.
problem Optimal waste collection in urban areas to manage increasing waste volumes.
method Predictive algorithms based on historical and future data for route planning.
result Reduction in distance traveled and operational costs by 70%.
Unified framework connects credit risk metrics with information theory.
problem Disconnection between industry-standard metrics and statistical theory.
method Unified information-theoretic framework, proving IV equals PSI, deriving standard errors, formalizing trade-off, automated binning with XGBoost.
result Unified framework connects IV and PSI, providing statistical foundation for metrics.
New methods reduce bias in estimating calibration error.
problem Reducing bias in estimating calibration error.
method Synthesizing model outputs and using equal-mass bins.
result Two reliable calibration-error estimators found: debiased estimator and ECE_sweep.
We propose a spatial diffuseness feature for deep neural network (DNN)-based automatic speech recognition to improve recognition accuracy in reverberant and noisy environments. The feature is computed in real-time from multiple microphone signals without requiring knowledge or estimation of the direction of arrival, an…
This paper improves multi-class calibration methods using mutual information maximization-based binning.
problem Calibration of deep neural network predictions, especially for small prior classes.
method I-Max concept for binning, shared class-wise calibration strategy.
result Improves multi-class ranking and calibration performance using a small calibration set.
Paper analyzes ECE bias and provides bounds for its estimation.
problem Understanding the estimation bias in ECE for machine learning models.
method Information-theoretic approach to analyze bias in uniform mass and uniform width binning strategies.
result Established upper bounds on ECE estimation bias and optimal number of bins.
New methods improve estimation of nonhomogeneous Poisson processes from limited data.
problem Estimating nonhomogeneous Poisson processes from limited data.
method Formulated as a learning generalization problem, proposed adaptive and data-driven binning methods.
result Improved estimation of nonhomogeneous Poisson processes with limited data.
A new method detects and displays pairwise dependence between variates.
problem Detecting and visualizing dependence between variates of different types.
method Recursive random binning with approximations to Pearson's statistic.
result The method is well-calibrated and powerful against common test alternatives.
ABM automates feature engineering and variable selection for loss-based models.
problem Improving model performance through better feature engineering and variable selection.
method ABM uses group and fused lasso regularization to automatically select cutting points and variables.
result ABM integrates feature engineering, variable selection, and model training.
Regularized contextual bandits use bins to solve multi-armed bandit problems.
problem Contextual bandit problems with a known baseline policy.
method Nonparametric model, splitting context space into bins, solving bandit instances independently.
result Intermediate convergence rates interpolating between slow and fast rates.
Recent advances in statistical theory, together with advances in the computational power of computers, provide alternative methods to do mass-univariate hypothesis testing in which a large number of univariate tests, can be properly used to compare MEEG data at a large number of time-frequency points and scalp location…
A method for non-parametric conditional distribution estimation using CRPS-optimal binning.
problem Non-parametric conditional distribution estimation.
method Partitioning covariate-sorted observations into bins to minimize LOO-CRPS, selecting K by K-fold cross-validation of test CRPS.
result Produces narrower prediction intervals with near-nominal coverage compared to split-conformal competitors.
Paper tackles flexible bin packing for e-commerce, reducing costs.
problem Optimizing packing of cuboid items into bins with minimal surface area.
method Multi-task Selected Learning approach to generate item packing sequence and orientation.
result Selected Learning method achieves 5.47% cost reduction compared to greedy algorithms.
Solves online 3D bin packing with deep reinforcement learning under constraints.
problem Challenges of packing items immediately without information and constraints.
method Constrained deep reinforcement learning (DRL) with feasibility predictor.
result Significantly outperforms state-of-the-art methods in online 3D bin packing.
Method infers causal direction using data discretization and complexity calculation.
problem Determining causal direction between continuous variables.
method MDL Binning technique for data discretization and complexity calculation.
result Captures the shape of the data to determine causal direction.
Probability Density Estimation (PDE) is a multivariate discrimination technique based on sampling signal and background densities defined by event samples from data or Monte-Carlo (MC) simulations in a multi-dimensional phase space. In this paper, we present a modification of the PDE method that uses a self-adapting bi…
This paper introduces minimum-risk recalibration for probabilistic classifiers, improving their reliability and accuracy.
problem Improving the reliability and accuracy of probabilistic classifiers.
method Minimum-risk recalibration within the MSE decomposition framework, analyzing UMB method and label shift adaptation.
result The optimal number of bins for UMB scales with n1/3, resulting in a risk bound of approximately O(n−2/3). A new survival analysis method eliminates hyperparameter tuning.
problem Survival analysis model hyperparameters are difficult to tune.
method Flexible survival density estimation with importance sampling.
result Matches or outperforms existing methods on real-world datasets.
Robust nonparametric regression method removes adversarial noise effectively.
problem Nonparametric regression under adversarial noise contamination.
method Local binning median followed by kernel smoothing or local polynomial regression.
result Minimax optimality over Hölder and Sobolev classes with arbitrary smoothness.
System separates sounds from mixtures without ground truth info.
problem Sound separation from multi-channel mixtures without labeled data.
method Deep clustering on multi-channel mixtures, projecting bins to spatially correlated clusters.
result Performance matches ground truth separation using only multi-channel mixtures.
This study examines how discretization improves neural forecasting models.
problem Improving predictive performance of neural forecasting models.
method Empirical investigation of data binning techniques on various neural forecasting architectures.
result Data binning almost always improves forecasting accuracy, but the type of binning is less important.
Estimates sample size for subgroup analysis in randomized experiments.
problem Determining sample size for accurate subgroup analysis.
method Turns inference problem into simultaneous inference, calculates sample size based on confidence level and margin of error.
result Allows inversion of sample size to feasible number of treatment arms or partition complexity.
Ranked Reward algorithm improves bin packing performance.
problem Improving reinforcement learning for combinatorial optimization.
method Ranking rewards from self-play to create a relative performance metric.
result Ranked Reward algorithm outperforms other methods on bin packing problems.
Paper tackles noisy, biased object detection in cluttered environments.
problem Noisy and biased object detection in unknown cluttered environments.
method Incremental active semi-supervised learning (IASSL) combining batch-based active learning and bin-based semi-supervised learning.
result Superior performance compared to state-of-the-art object detection methods.
New method unfolds distribution moments directly from data without binning.
problem Deconvolving detector distortions in particle physics.
method Uses machine learning, inspired by GANs, to unfold moments directly.
result More precise than bin-based approaches and comparable to unbinned methods.
In this paper we perform a statistical analysis over the returns and relative prices of the CAC 40 and the S\&P 500 with the purpose of analyzing the intra-day seasonalities of single and cross-sectional stock dynamics. In order to do that, we characterized the dynamics of a stock (or a set of stocks) by the evolut…
A new DP algorithm improves privacy in hashing and sampling for search and learning.
problem Improving privacy in hashing and sampling for large-scale applications.
method Combines differential privacy with one permutation hashing and bin-wise consistent weighted sampling.
result Proposes DP-OPH and DP-BCWS algorithms that enhance privacy while maintaining utility.
MCLNN improves sound recognition by learning frequency bands.
problem Sound recognition from neural networks often misses environmental sound specifics.
method MCLNN incorporates filterbank behavior and automates feature combination exploration.
result MCLNN outperforms state-of-the-art methods on ESC-10 dataset.
New method improves model calibration efficiency and accuracy.
problem Improving model calibration for better probability estimates.
method Scaling-binning calibrator method that reduces variance and ensures calibration.
result 35% lower calibration error than histogram binning and guarantees true calibration.
BIN and CBN models infer health variables from symptoms and signals.
problem Inferring health variables from symptoms and signals.
method Bidirectional inference networks (BIN) and composite BIN (CBIN).
result CBIN achieves state-of-the-art performance and better accuracy.
Efficiently models complex spatiotemporal patterns with large datasets.
problem Flexible modeling of large-scale spatiotemporal point patterns.
method Cox process inference using Fourier features for scalable inference.
result Approximate Bayesian method fits over 100,000 events in 3D.
PoET-BiN reduces power consumption in neural networks on embedded devices.
problem Power inefficiency in neural network implementations on embedded platforms.
method Look-Up Table based implementation with a modified Decision Tree approach.
result Near state-of-the-art results with up to 6 orders of magnitude energy reduction.