A new method for causal inference in high-dimensional data using machine learning.
problem Causal inference in high-dimensional observational data.
method Support Points Sample Splitting (SPSS) for efficient double machine learning (DML) in causal inference.
result Deep learning with SPSS and hybrid methods outperform SVM with SPSS in computational efficiency and estimation quality.
SPlit optimizes dataset splitting for better model performance.
problem Improving model performance through optimal dataset splitting.
method Adapting Support Points (SP) algorithm for subsampling and categorical variables in a sequential nearest neighbor approach.
result SPlit significantly improves worst-case testing performance compared to random splitting.
The paper proves conditions for non-uniform expansion in partially hyperbolic systems.
problem Conditions for non-uniform expansion in partially hyperbolic systems.
method Analysis of Lyapunov exponents and dominated splittings.
result Existence of physical SRB measure under specific conditions.
SBSS uses similarity to split data for better classifier training.
problem Training better classifiers with realistic performance estimation.
method SBSS uses both input and output space information to split data using similarity functions.
result SBSS outperformed ordinary stratified 10-fold cross-validation in 75% of scenarios.
MaxRR efficiently unlearns models by splitting and selecting core samples.
problem Efficiently unlearning models while verifying unlearning guarantees.
method Model splitting and core sample selection with a generalized unlearning metric.
result MaxRR achieves efficient unlearning with properties matching full retraining.
A method to split a data point into two parts that individually cannot reconstruct the whole, but together can.
problem Splitting a single data point into two parts such that neither can reconstruct the whole but together can.
method Borrowing ideas from Bayesian inference to achieve a continuous analog of data splitting.
result A method to achieve data fission, enabling post-selection inference in finite samples.
Generative model uses random weighted support points for interpretable data sampling.
problem Creating diverse and interpretable sample sets from large datasets efficiently.
method Random weighted support points from Dirichlet process and Bayesian bootstrap.
result High-quality and diverse outputs at lower computational cost.
This work improves data reconstruction methods by ensuring unique solutions and refining optimization.
problem Ensuring unique solutions and optimizing reconstruction from KKT conditions.
method Discussion of sufficient conditions for unique solutions and introduction of sample splitting for optimization.
result Sample splitting improves reconstruction performance across various methods.
Paper introduces a new sampling method for Bayesian inference.
problem Efficiently sampling from complex posterior distributions.
method Plug-and-Play split Gibbs sampler using variable splitting and ADMM.
result The method allows for embedding deep generative priors in Bayesian inference.
New approach makes survival analysis fairer without specifying sensitive features.
problem Ensuring fairness in survival analysis models across different subpopulations.
method Distributionally robust optimization (DRO) with sample splitting strategy.
result Converted existing survival analysis models into fair versions without specifying sensitive features.
This work approximates full conformal prediction for neural networks without sample splitting.
problem Uncertainty quantification for neural network regression models.
method Approximating full conformal prediction using Gauss-Newton influence for post-hoc uncertainty estimation.
result Locally-adaptive and often tighter prediction intervals compared to split-CP.
Paper proposes a novel SVM method for creating survival trees.
problem Creating non-linear survival trees for right-censored data.
method L2-regularized dipole splitting criteria with kernel methods.
result Non-linear splits using polynomial and Gaussian kernels show similar predictive power but often smaller tree sizes.
MCP extends conformal prediction to vector-valued score functions without data splitting.
problem Fixed prediction set shapes in scalar score functions limit coverage guarantees.
method MCP uses a single optimization problem for prediction set design and calibration, eliminating data splitting.
result RemMCP and RelMCP achieve target coverage with smaller or comparable prediction set sizes, reducing variance.
Study robustness of split conformal prediction under adversarial attacks.
problem Ensuring distribution-free coverage guarantees in CP under adversarial conditions.
method Theoretical analysis and extensive experiments on split conformal prediction robustness.
result Prediction coverage varies with calibration-time attack strength, enabling control over coverage under adversarial tests.
A new method estimates corporate bond defaults in financial networks efficiently.
problem Challenges in valuing corporate bonds in interconnected financial systems.
method Bi-Level Importance Sampling with Splitting
result The method efficiently estimates rare default events in financial networks.
The paper splits manifolds using infinity harmonic functions with linear growth.
problem Splitting manifolds with specific harmonic functions.
method Analyzes manifolds with non-negative Ricci or sectional curvature, focusing on infinity harmonic functions with linear growth.
result Extends Savin's theorem to surfaces with non-negative sectional curvature.
Performing exact Bayesian inference for complex models is computationally intractable. Markov chain Monte Carlo (MCMC) algorithms can provide reliable approximations of the posterior distribution but are expensive for large datasets and high-dimensional models. A standard approach to mitigate this complexity consists i…
We provide necessary conditions for the Alexander polynomials of algebraically split component-preservingly amphicheiral links. We raise a conjecture that the Alexander polynomial of an algebraically split component-preservingly amphicheiral link with even components is zero. Our necessary conditions and some examples …
It is widely believed that the prediction accuracy of decision tree models is invariant under any strictly monotone transformation of the individual predictor variables. However, this statement may be false when predicting new observations with values that were not seen in the training-set and are close to the location…
Adaptive Multilevel Splitting improves rare event pricing for financial derivatives.
problem Efficient pricing of binary options in rare event regimes with discontinuous payoffs.
method Adaptive Multilevel Splitting (AMS) reformulates rare-event problem as conditional events.
result AMS achieves up to 200-fold improvements over standard Monte Carlo, preserving unbiasedness.
A new method speeds up sampling in diffusion models.
problem Slow sample generation in diffusion models.
method Proposed Splitting Integrators for fast stochastic sampling.
result Achieved FID score of 2.36 in 100 NFE, significantly faster than baselines.
Develops a new inference method for split-sample estimators using multiple splits.
problem Statistical dependence and variability in split-sample estimators.
method Averaging across multiple splits, proving a central limit theorem, and developing new inference approaches.
result Valid confidence intervals and improved power in comparing model performance.
New method for estimating out-of-sample R² from gene expression data.
problem Lack of a well-defined and unbiased estimator for out-of-sample R².
method Explicitly defined out-of-sample R², provided an unbiased estimator, and calculated standard error.
result Demonstrated improved model comparison for gene expression phenotypes.
In this paper we propose a fast online Kernel SVM algorithm under tight budget constraints. We propose to split the input space using LVQ and train a Kernel SVM in each cluster. To allow for online training, we propose to limit the size of the support vector set of each cluster using different strategies. We show in th…
FastForest boosts Random Forest speed by 24%.
problem Efficiency in processing speed for Random Forest.
method Subsample Aggregating, Logarithmic Split-Point Sampling, Dynamic Restricted Subspacing.
result Average 24% increase in processing speed with accuracy maintained.
Estimator calculates surface curvature from point cloud samples.
problem Accurately estimating curvature from limited point cloud data.
method Algorithm using probability distribution and nearby points control.
result Controlled number of points ensures accurate curvature estimation.
Motivation: With the development of droplet based systems, massive single cell transcriptome data has become available, which enables analysis of cellular and molecular processes at single cell resolution and is instrumental to understanding many biological processes. While state-of-the-art clustering methods have been…
New approach connects stochastic gradient descent to ODE splitting schemes.
problem Improving convergence in stochastic optimization.
method Connection between stochastic gradient descent and ODE splitting schemes.
result Derive a new upper bound on global splitting error.
In this short note we observe that the sample complexity of PAC machine learning of various concepts, including learning the maximum (EMX), can be exactly determined when the support of the probability measures considered as models satisfies an a-priori bound. This result contrasts with the recently discovered undecida…
The paper proves consistency of archetypal analysis for multivariate data.
problem Finding optimal archetype points for multivariate data.
method Uses convex polytope to summarize data, proving consistency under specific distribution assumptions.
result Archetype points converge to optimal solution under certain conditions.
A new method for kernel tests without data splitting increases power.
problem Lack of power in kernel-based tests due to data splitting.
method Selective inference framework to learn hyperparameters and test on full sample.
result Empirically larger test power without data splitting, regardless of split proportion.
Each compact Riemannian manifold with no conjugate points admits a family of functions whose integrals vanish exactly when central Busemann functions split linearly. These functions vanish when all central Busemann functions are sub- or superharmonic. When central Busemann functions are convex or concave, they must be …
The paper studies splitting maps in link Floer homology using skein exact sequences.
problem Understanding splitting maps for links using link Floer homology.
method Link Floer homology and skein exact sequences.
result Splitting maps for torus links T(n,n) are associated with integer points in permutahedra. New framework for regression trees with multivariate response and dynamic mean vectors.
problem Characterizing and implementing regression trees for multivariate responses.
method High dimensional model with dynamic mean vectors over multi-dimensional change axes.
result Optimal rate of convergence and asymptotic valid confidence intervals for change points.
New proof of Giroux Correspondence for tight contact 3-manifolds.
problem Proving the Giroux Correspondence for tight contact 3-manifolds.
method Introducing tight Heegaard splittings, using refinement process, and translating moves between splittings to moves between open books.
result Proves the tight Giroux Correspondence for contact 3-manifolds.
The paper explores how splitting data samples influences optimal neural network hyperparameters.
problem Understanding the effectiveness of neural networks and their hyperparameters.
method Investigates the role of sample splitting in neural network hyperparameter selection.
result Optimal hyperparameters derived from sample splitting lead to a neural network model that minimizes prediction risk asymptotically.
ICP improves prediction intervals for continuous outcomes at lower computational cost.
problem Systematic bias in point predictions that undermines their use in decision-making.
method Develops Isotonic Conformal Prediction (ICP) framework to decouple calibration from prediction-set construction.
result SICP and TICP procedures match SC-CP coverage at lower computational cost.
Optimally estimates a functional using nuisance function tuning and sample splitting.
problem Estimating optimal rates for a doubly robust functional.
method Combines nuisance function tuning and sample splitting strategies.
result Shows optimal rates of convergence for various estimators.
New privacy-preserving method for conformal prediction without splitting data.
problem Privacy and uncertainty quantification in data-driven decision making.
method Proposes a full-data privacy-preserving conformal prediction framework using differential privacy.
result Demonstrates improved prediction sets compared to split-based private baselines.
Constructs ε-splitting maps for geodesic balls with non-negative Ricci curvature.
problem Constructing ε-splitting maps for geodesic balls with non-negative Ricci curvature.
method Induction and stratified almost Gou-Gu Theorem for finding directional points; error estimates for projections.
result Constructs ε-splitting maps on concentric geodesic balls with uniformly small radius. This work introduces a transformation-based learner model for classification forests. The weak learner at each split node plays a crucial role in a classification tree. We propose to optimize the splitting objective by learning a linear transformation on subspaces using nuclear norm as the optimization criteria. The le…
Invariant counts maximum stable umbilic splits.
problem Counting stable umbilic splits on surfaces.
method Introducing an invariant to count maximum stable umbilic splits.
result Established properties of the multiplicity invariant.
This paper provides estimation and inference methods for an identified set's boundary (i.e., support function) where the selection among a very large number of covariates is based on modern regularized tools. I characterize the boundary using a semiparametric moment equation. Combining Neyman-orthogonality and sample s…
Exact distribution of split conformal prediction coverage found.
problem Determining the reliability of prediction sets in batch mode.
method Analysis of exchangeable data to find universal distribution of empirical coverage.
result Exact distribution of empirical coverage is universal and determined by nominal miscoverage level and calibration sample size.
New algorithms improve sampling from constrained distributions.
problem Generating samples from distributions under constraints.
method Kinetic Langevin dynamics and splitting schemes.
result Improved complexity bounds over existing methods.
This paper introduces a class of k-nearest neighbor (k-NN) estimators called bipartite plug-in (BPI) estimators for estimating integrals of non-linear functions of a probability density, such as Shannon entropy and Rényi entropy. The density is assumed to be smooth, have bounded support, and be uniformly bounded from…
Optimizes data splitting for shorter conformal prediction intervals.
problem Minimizing prediction interval length while maintaining coverage.
method Theoretical framework for optimal data splitting in split conformal prediction.
result Analytical characterizations of length-optimal split ratios in various settings.
The paper studies statistical properties of CART regression trees.
problem Understanding the statistical properties of CART regression trees.
method The paper constructs a prior distribution on split points and solves a nonlinear optimization problem to bound the Pearson correlation between the optimal decision stump and response data.
result CART with cost-complexity pruning achieves an optimal complexity/goodness-of-fit tradeoff when the depth scales with the logarithm of the sample size.