Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

51102153204 · May 202619922001200920172026
48 results for split scoring

MCP extends conformal prediction to vector-valued score functions without data splitting.

problem Fixed prediction set shapes in scalar score functions limit coverage guarantees.
method MCP uses a single optimization problem for prediction set design and calibration, eliminating data splitting.
result RemMCP and RelMCP achieve target coverage with smaller or comparable prediction set sizes, reducing variance.

We propose an algorithm named best-scored random forest for binary classification problems. The terminology "best-scored" means to select the one with the best empirical performance out of a certain number of purely random tree candidates as each single tree in the forest. In this way, the resulting forest can be more …

2019-05-27abs ↗pdf ↗

Improves robustness of propensity score estimators in challenging settings.

problem Limited overlap, small sample sizes, or unbalanced data.
method Extends calibration techniques for propensity score models, focusing on sample-splitting schemes.
result Calibration reduces variance and bias in inverse probability weighting and double/debiased machine learning frameworks.

Proposes a new method to minimize non-singleton predictions in conformal prediction.

problem Large prediction sets in conformal prediction are costly and inefficient.
method Introduces a new nonconformity score to minimize non-singleton sets and provides an algorithm to compute it efficiently.
result The proposed Singleton-Optimized Conformal Prediction (SOCOP) method increases singleton frequency by over 20% compared to standard scores, with minimal impact on average set size.

Paper proposes a new method for conditional coverage in conformal prediction.

problem Lack of strong conditional coverage guarantees in existing conformal prediction methods.
method Modified non-conformity score using local approximation of conditional distribution.
result Unified framework and empirical evaluations show advantage of the new method.

New privacy-preserving method for conformal prediction without splitting data.

problem Privacy and uncertainty quantification in data-driven decision making.
method Proposes a full-data privacy-preserving conformal prediction framework using differential privacy.
result Demonstrates improved prediction sets compared to split-based private baselines.

Skew-adaptive method improves prediction intervals for regression.

problem Improving prediction intervals for regression models, especially in cases of skewness and varying scales.
method Develops a skew-adaptive extension of split conformal prediction using an asymmetric interval family and gauge approach.
result Preserves marginal validity and adapts to local scale and skewness, with efficiency gains over existing methods.

Study robustness of split conformal prediction in data contamination setting.

problem Robustness of split conformal prediction under data contamination.
method Analyze split conformal prediction's performance in a contaminated data setting and propose a new method.
result Demonstrated the impact of corrupted data on prediction intervals' coverage and efficiency.

LoBoost improves local conformal prediction for gradient-boosted trees without extra data splits.

problem Quantifying uncertainty in gradient-boosted tree predictions.
method Model-native local conformal prediction using leaf structure.
result Competitive interval quality and improved test MSE with large calibration speedups.

New approach makes survival analysis fairer without specifying sensitive features.

problem Ensuring fairness in survival analysis models across different subpopulations.
method Distributionally robust optimization (DRO) with sample splitting strategy.
result Converted existing survival analysis models into fair versions without specifying sensitive features.

This work improves SINDy-type algorithms for system identification using score-guided dictionary selection.

problem Improving accuracy and interpretability in dynamical system identification.
method Score-guided library selection to refine dictionary terms in sparse regression.
result Score-guided methods enhance SINDy's robustness in discovering governing equations.

This work approximates full conformal prediction for neural networks without sample splitting.

problem Uncertainty quantification for neural network regression models.
method Approximating full conformal prediction using Gauss-Newton influence for post-hoc uncertainty estimation.
result Locally-adaptive and often tighter prediction intervals compared to split-CP.

New sampling technique improves boosting model accuracy.

problem Improving generalization performance and learning time in stochastic gradient boosting.
method Formulated optimization problem to maximize estimation accuracy, leading to Minimal Variance Sampling (MVS).
result MVS significantly increases model quality and reduces the number of examples needed.

The paper decomposes probabilistic scores into reliability, uncertainty, and information loss.

problem Understanding the reliability and uncertainty of probabilistic predictions.
method Developed decomposition identities for proper losses, quantifying reliability, residual uncertainty, and information gain.
result A three-term identity for classification scores, revealing miscalibration, grouping term, and feature-level uncertainty.

Proposes oblique predictive clustering trees for faster, more efficient learning.

problem Learning time scales poorly with output space dimensionality and cannot exploit data sparsity.
method Design and implement oblique splits using linear combinations of features.
result Achieves performance on-par with state-of-the-art methods and is orders of magnitude faster.

Develops methods to simulate rare transitions in molecular systems.

problem Rare transitions between metastable states in molecular systems are difficult to study due to limited data.
method Two novel methods: chain-based and midpoint-based approaches.
result Demonstrates effectiveness of methods in both data-rich and data-scarce scenarios.

Improved flow-based models capture dependencies better with multi-scale autoregressive priors.

problem Limited expressiveness of flow-based models for long-range data dependencies.
method Introducing channel-wise dependencies through multi-scale autoregressive priors (mAR) in split coupling flow layers (mAR-SCF).
result Achieves state-of-the-art density estimation results on MNIST, CIFAR-10, and ImageNet.

The Split-Session Cluster GARCH model captures tail heterogeneity in overnight and intraday returns.

problem Capturing tail behavior and dependence in multivariate asset returns.
method Convolution-tt distributions, session and sector clustering, block-structured correlation matrices.
result Session-specific and sector-level tail parameters improve model fit and out-of-sample performance.

Enhanced conformal methods improve validity of LLM outputs.

problem Lack of conditional validity and high false rejection rates in LLM validity guarantees.
method Adaptive conditional conformal procedure and improved scoring function differentiation.
result Demonstrated improved validity and utility on real datasets.

Study shows pooling scores for conformal prediction distorts group coverage.

problem Pooling scores for conformal prediction distorts group coverage.
method Derived conservation law and lower bound, demonstrated tension between fairness definitions, quantified trade-off between policies.
result Pooling scores for conformal prediction distorts group coverage.

Students hiring ghostwriters to write their assignments is an increasing problem in educational institutions all over the world, with companies selling these services as a product. In this work, we develop automatic techniques with special focus on detecting such ghostwriting in high school assignments. This is done by…

2019-06-04abs ↗pdf ↗

Improved multivariate conformal prediction by standardizing residuals.

problem Weak conditional coverage in heteroskedastic multivariate settings.
method Natural extension of univariate normalization to multivariate setting, whitening residuals and standardizing local variance.
result Standardized residuals yield asymptotic conditional coverage under certain distributions.

GBOC detects anomalies in time series data using granular-ball vectors.

problem Challenges in modeling normal behavior in dynamic, nonlinear time series data.
method Granular-ball Vector Data Description (GVDD) and Granular-ball One-Class Network (GBOC).
result GBOC improves anomaly detection in time series data.

CCI combines Bayesian and gradient boosting to create fair, reliable credit risk scores.

problem Tackles high-stakes lending decisions with changing data distributions and fairness constraints.
method Combines Bayesian neural risk scorer and fairness-constrained gradient boosting with shift-aware fusion.
result CCI achieves best trade-off between discrimination, calibration, stability, and fairness.

A method for non-parametric conditional distribution estimation using CRPS-optimal binning.

problem Non-parametric conditional distribution estimation.
method Partitioning covariate-sorted observations into bins to minimize LOO-CRPS, selecting K by K-fold cross-validation of test CRPS.
result Produces narrower prediction intervals with near-nominal coverage compared to split-conformal competitors.

TRACE improves conformal prediction for multi-dimensional outputs.

problem Challenges in constructing valid and informative conformal prediction regions for multi-dimensional outputs.
method TRACE uses transport alignment in diffusion and flow matching models to define nonconformity scores.
result TRACE yields valid and adaptive conformal prediction regions for multimodal and non-convex distributions.

This work introduces COLA, a strategy to aggregate conformal prediction sets efficiently.

problem Efficiently combining multiple conformity scores to reduce prediction set size.
method Introduces COnfidence-Level Allocation (COLA) to optimally allocate confidence levels across sets.
result COLA achieves smaller prediction sets than state-of-the-art methods while maintaining valid coverage.

GT-Score reduces overfitting in trading strategies by integrating multiple criteria.

problem Overfitting in data-driven financial models leads to unreliable out-of-sample performance.
method Integrates performance, statistical significance, consistency, and downside risk into a composite objective function.
result Improves generalization ratio by 98% compared to baseline objective functions in walk-forward validation.

Develops conformalized prediction intervals for bounded continuous outcomes.

problem Predicting continuous outcomes within bounded ranges, especially when models are misspecified.
method Conformal prediction intervals based on transformation regression models, accounting for heteroscedasticity and asymmetry.
result Valid finite-sample coverage confirmed in simulations and real data applications.

Paper develops a new objective for hierarchical clustering in Euclidean space.

problem Hierarchical clustering in Euclidean space with dissimilarity scores.
method Develops a new global objective and connects it to bisecting k-means.
result Optimal 2-means solution approximates the new objective, proving bisecting k-means optimizes a natural global objective.

SketchBoost accelerates GBDT for multioutput problems up to 40x.

problem Efficiently training GBDT for multioutput problems with high-dimensional outputs.
method Approximate computation of scoring function for faster decision tree splitting.
result SketchBoost speeds up GBDT training by up to 40 times.

Automated multi-task learning algorithm that optimizes network topology.

problem Over-sharing in multi-task learning leads to over-generalization and suboptimal performance.
method Tree-structured design space with gumbel-softmax sampling for differentiable network splitting.
result End-to-end trainable algorithm that optimizes network topology for multiple objectives across tasks.

We propose a novel method designed for large-scale regression problems, namely the two-stage best-scored random forest (TBRF). "Best-scored" means to select one regression tree with the best empirical performance out of a certain number of purely random regression tree candidates, and "two-stage" means to divide the or…

2019-05-09abs ↗pdf ↗

Novel Bayesian framework for Poisson inverse problems using Bregman geometry.

problem Solving Poisson inverse problems with non-Euclidean geometry and positivity constraints.
method Develops a Monte Carlo sampling algorithm that accounts for Bregman geometry, data augmentations, and conditional conjugacy properties.
result Efficient sampling via Gibbs steps and Hessian Riemannian Langevin Monte Carlo (HRLMC) for positivity constraints.

A novel bootstrap method improves concept drift detection in predictive models.

problem Detecting changes in predictive relationships (concept drift) in data-driven applications.
method Developed a nested bootstrap procedure to calibrate control limits using the entire initial sample.
result The method yields more accurate baseline models and faster CL setup times.

Unified framework for testing deep learning models with concept activation vectors.

problem Statistical instability and discontinuity in testing with concept activation vectors.
method Introducing α-TCAV, a generalized framework that replaces the indicator function with a parameterized smooth function.
result Unified probabilistic formulation that subsumes TCAV and Multi-TCAV, providing principled guidance on tuning the parameter.