Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,878 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Nov 199319922001200920172026
48 results for distributional splitting

New random forest criteria improve splitting for non-location structured data.

problem Improving random forest splitting for non-location structured data.
method Implement and compare various distributional splitting criteria inside a single honest-forest implementation.
result Distributional splitting criteria, especially sliced-Wasserstein, improve performance on multivariate responses.

Exact distribution of split conformal prediction coverage found.

problem Determining the reliability of prediction sets in batch mode.
method Analysis of exchangeable data to find universal distribution of empirical coverage.
result Exact distribution of empirical coverage is universal and determined by nominal miscoverage level and calibration sample size.

To reduce the label complexity in Agnostic Active Learning (A^2 algorithm), volume-splitting splits the hypothesis edges to reduce the Vapnik-Chervonenkis (VC) dimension in version space. However, the effectiveness of volume-splitting critically depends on the initial hypothesis and this problem is also known as target…

2018-09-28abs ↗pdf ↗

Combining Bayesian deep learning and split conformal prediction affects out-of-distribution coverage.

problem Improving out-of-distribution coverage in multiclass image classification.
method Combining Bayesian deep learning with split conformal prediction methods.
result Combining methods can reduce out-of-distribution coverage in some cases.

Novel methods for splitting Gaussian mixtures improve uncertainty propagation in nonlinear systems.

problem Improving accuracy and efficiency in nonlinear uncertainty propagation.
method Preserving mean and covariance, novel heuristics for selecting splitting direction informed by initial uncertainty and nonlinear function properties.
result Improved accuracy and efficiency in uncertainty propagation compared to existing techniques.

A new type of distributional regression tree uses soft split rules for better predictive performance.

problem Estimating complete conditional distributions in regression.
method Distributional adaptive soft regression trees using multivariate soft split rules.
result The method outperforms various benchmark methods, especially in complex non-linear interactions.

Split conformal prediction provides finite-sample guarantees for black-box models without distributional assumptions.

problem Weak performance guarantees for modern predictive models under minimal assumptions.
method Develops finite-sample guarantees for split conformal prediction, a method that uses nested prediction sets and order statistics.
result The coverage of prediction sets based on order statistics stochastically dominates the Beta distribution.

SPlit optimizes dataset splitting for better model performance.

problem Improving model performance through optimal dataset splitting.
method Adapting Support Points (SP) algorithm for subsampling and categorical variables in a sequential nearest neighbor approach.
result SPlit significantly improves worst-case testing performance compared to random splitting.

A method to split a data point into two parts that individually cannot reconstruct the whole, but together can.

problem Splitting a single data point into two parts such that neither can reconstruct the whole but together can.
method Borrowing ideas from Bayesian inference to achieve a continuous analog of data splitting.
result A method to achieve data fission, enabling post-selection inference in finite samples.

In-network learning outperforms Federated and Split learning in wireless networks.

problem Efficiently using distributed features for inference in wireless networks.
method Proposes 'in-network learning' architecture, uses neural networks for optimization, compares with Federated and Split learning.
result In-network learning offers better accuracy and bandwidth savings.

Split learning improves deep learning in healthcare by sharing data.

problem Data scarcity in healthcare limits deep learning applications.
method Distributed learning approach for collaborative training of deep neural networks.
result Split learning outperformed non-collaborative methods in both binary and multi-label classification tasks.

This paper compares communication efficiency of split learning and federated learning in various scenarios.

problem Comparing communication efficiency of split learning and federated learning in different settings.
method Examined various practical scenarios of distributed learning setups and compared the two methods.
result Communication efficiency of split learning and federated learning depends on the number of clients, model size, and data samples.

Data thinning splits observations into independent parts for convolution-closed distributions.

problem Validation of unsupervised learning results in settings with limited data.
method Data thinning, splitting observations into independent parts following the same distribution.
result Data thinning provides an attractive alternative to cross-validation in settings with limited sample splitting.

The paper proves integral formulas for manifolds with multiple orthogonal distributions.

problem Understanding geometric properties of manifolds with multiple orthogonal distributions.
method Develops integral formulas for Riemannian manifolds with k>2k>2 orthogonal complementary distributions.
result Generalizes known formulas for k=2k=2 and applies to manifold splitting and immersions.

In this paper, we discuss an extension of the Split Hamiltonian Monte Carlo (Split HMC) method for Gaussian process model (GPM). This method is based on splitting the Hamiltonian in a way that allows much of the movement around the state space to be done at low computational cost. To this end, we approximate the negati…

2012-01-19abs ↗pdf ↗

Develops significance tests for neural networks without strong assumptions or excessive computation.

problem Addressing the black-box nature of deep neural networks for feature relevance testing.
method Derives one-split and two-split tests relaxing assumptions and computational complexity.
result Establishes asymptotic null distributions and consistency in Type II error.

SVHN dataset's split affects generative models but not digit classification.

problem Distribution mismatch between SVHN training and test sets impacts generative models.
method Empirically showed distribution mismatch affects generative models; proposed mixing and re-splitting.
result Distribution mismatch in SVHN dataset significantly impacts probabilistic generative models.

New classification of conformal structures with maximal G2G_2 symmetry.

problem Classifying conformal structures with maximal G2G_2 symmetry.
method Complete local classification of homogeneous 4D split-conformal structures.
result Established a complete local classification of conformal structures with maximal G2G_2 symmetry.

Combines neural networks with splitting-up method for filtering equations.

problem Approximating the solution of filtering equations for signal processes.
method Combines splitting-up method with neural networks.
result Produces an approximation of the unnormalised conditional distribution.

New algebraic structures for Lie 2-algebroids and their connections.

problem Characterizing and understanding Lie 2-algebroids and their structures.
method Construction of homotopy Poisson algebra and introduction of Dirac structures.
result One-to-one correspondence between Manin triples and Lie 2-bialgebroids.

Optimizes data splitting for shorter conformal prediction intervals.

problem Minimizing prediction interval length while maintaining coverage.
method Theoretical framework for optimal data splitting in split conformal prediction.
result Analytical characterizations of length-optimal split ratios in various settings.

A random Heegaard splitting is a 3-manifold obtained by using a random walk of length n on the mapping class group as the gluing map between two handlebodies. We show that the joint distribution of random walks of length n and their inverses is asymptotically independent, and converges to the product of the harmonic an…

2008-09-29abs ↗pdf ↗

Sparse covariance estimation in the vertical-split model achieves exponential improvement over dense estimates.

problem Minimax estimation error for distributed covariance matrix estimation in the vertical-split setting.
method Elementwise ss-sparsity is shown to reduce communication and sample complexity.
result Minimax lower bounds for 11-sparse cross-covariance estimation are established.

Study shows splitting schemes can approximate WFR flows faster than the exact flow.

problem Improving sampling efficiency in Wasserstein-Fisher-Rao gradient flows.
method Investigates operator splitting techniques to numerically approximate WFR flows.
result A judicious choice of step size and operator ordering can lead to faster convergence of split schemes to the target distribution.

ES-MLP combines Graph-MLP with edge splitting for node classification on both homophilic and heterophilic graphs.

problem Node classification on graphs with mixed homophilic and heterophilic properties.
method Combines Graph-MLP with edge splitting mechanism from ES-GNN to learn two adjacency matrices based on relevant and irrelevant feature pairs.
result ES-MLP achieves performance comparable to homophilic and heterophilic models without using edges during inference.

Study robustness of split conformal prediction under adversarial attacks.

problem Ensuring distribution-free coverage guarantees in CP under adversarial conditions.
method Theoretical analysis and extensive experiments on split conformal prediction robustness.
result Prediction coverage varies with calibration-time attack strength, enabling control over coverage under adversarial tests.

Study cone structures on contact manifolds to understand their geometric properties.

problem Characterize cone structures on holomorphic contact manifolds.
method Characterize subadjoint varieties among Legendrian submanifolds in terms of contact prolongations.
result Holomorphic horizontal splitting of the canonical distribution on contact G-structures.

A new method for causal inference in high-dimensional data using machine learning.

problem Causal inference in high-dimensional observational data.
method Support Points Sample Splitting (SPSS) for efficient double machine learning (DML) in causal inference.
result Deep learning with SPSS and hybrid methods outperform SVM with SPSS in computational efficiency and estimation quality.

A new tree model, GRST, improves option pricing without log-normality assumptions.

problem Limitations of CRR binomial trees in valuing securities with early exercise characteristics.
method Gaussian Recombining Split Tree (GRST) that generates a discrete probability mass function approximating a Gaussian distribution.
result Option prices from GRST align closely with market prices.

Data splitting enhances model performance in overparametrized ridgeless regression.

problem Computational inefficiency in training models with large datasets.
method Data splitting as a regularization technique in overparametrized ridgeless regression.
result Data splitting improves statistical performance and computational complexity.

Study maximally symmetric distribution of An-Nurowski surface rolling on a plane.

problem Maximally symmetric (2,3,5)(2,3,5)-distribution of An-Nurowski surface rolling without slipping or twisting.
method Calculated vector fields defining a split g2\frak{g}_2 Lie algebra and projected to an action of SL(3,R)SL(3,\mathbb{R}).
result Obtained an action of SL(3,R)SL(3,\mathbb{R}) on the configuration space without a surface.

Batch-splitting (data-parallelism) is the dominant distributed Deep Neural Network (DNN) training strategy, due to its universal applicability and its amenability to Single-Program-Multiple-Data (SPMD) programming. However, batch-splitting suffers from problems including the inability to train very large models (due to…

2018-11-05abs ↗pdf ↗

Histogram binning method proven with guarantees without splitting data.

problem Proving theoretical guarantees for histogram binning without sample splitting.
method Using Markov property of order statistics to prove calibration guarantees for original method.
result Proves histogram binning has strong calibration guarantees without sample splitting.

Study robustness of split conformal prediction in data contamination setting.

problem Robustness of split conformal prediction under data contamination.
method Analyze split conformal prediction's performance in a contaminated data setting and propose a new method.
result Demonstrated the impact of corrupted data on prediction intervals' coverage and efficiency.