The paper outlines future work in random sets theory.
problem Developing a theory of statistical reasoning with random sets.
method Generalizing logistic regression, probability laws, and geometric uncertainty.
result A new geometric approach to uncertainty with general random sets.
We define two non-linear operations with random (not necessarily closed) sets in Banach space: the conditional core and the conditional convex hull. While the first is sublinear, the second one is superlinear (in the reverse set inclusion ordering). Furthermore, we introduce the generalised conditional expectation of r…
Introduces epistemic deep learning for better uncertainty estimation in neural networks.
problem Uncertainty quantification in deep neural networks.
method Random-set convolutional neural networks with belief function-based loss functions.
result Epistemic approach produces better performance in uncertainty estimation.
Study finds root vertex in large networks with high probability.
problem Finding the root vertex in large growing networks.
method Constructs confidence sets for the root vertex in various random network models.
result Confidence sets of size independent of the number of vertices contain the root vertex with high probability.
We determine the expected curvature polynomial of random real projective varieties given as the zero set of independent random polynomials with Gaussian distribution, whose distribution is invariant under the action of the orthogonal group. In particular, the expected Euler characteristic of such random real projective…
Random walks on Fuchsian Schottky groups have harmonic measures with lower dimension.
problem Understanding the dimensionality of harmonic measures for random walks.
method Analyzing finite range random walks on Fuchsian Schottky groups.
result Harmonic measures have dimension strictly less than the limit set's Hausdorff dimension.
Linear statistics of random zero sets are integrals of smooth differential forms over the zero set and as such are smooth analogues of the volume of the random zero set inside a fixed domain. We derive an asymptotic expansion for the variance of linear statistics of the zero divisors of random holomorphic sections of p…
Random subgroups in hyperbolic spaces have full limit sets and bounded critical exponents.
problem Understanding stationary random subgroups in hyperbolic spaces.
method Analyzing limit sets and critical exponents of random subgroups.
result Random subgroups have full limit sets and bounded critical exponents.
Study on spin random fields using chaos decomposition for cosmic microwave background modeling.
problem Modeling polarization of Cosmic Microwave Background using spin random fields.
method Explicit Wiener-Itô chaos decomposition of area measures of level sets.
result Reveals a clear difference between high frequency regime and zero spin case.
Random forests reduce bias and variance, especially in low SNR settings.
problem Reducing bias and variance in machine learning models, particularly in low SNR scenarios.
method Empirical study of random forests and bagging ensembles, focusing on the importance of mtry tuning. result Random forests reduce both bias and variance, outperforming bagging ensembles in high SNR settings.
Projected random forests improve circular data prediction with adaptive arc length and finite-sample coverage.
problem Regression with circular responses.
method Adapting linear-response models to circular data using projection and random forest out-of-bag mechanism.
result Projected random forest out-of-bag conformal prediction sets are more efficient and shorter than alternative methods.
New gradient coding schemes reduce decoding error in both random and adversarial straggler settings.
problem Creating efficient approximate gradient coding schemes for distributed optimization.
method Introduced novel approximate gradient codes based on expander graphs, achieving optimal decoding coefficients.
result Achieved nearly optimal error in random setting and nearly half the error in adversarial setting compared to existing codes.
New framework models epistemic uncertainty in GNNs using random sets.
problem Uncertainty in graph neural network predictions.
method Introduces a belief function (random set) approach to model epistemic uncertainty in GNNs.
result Demonstrates superior uncertainty quantification on various graph datasets.
Random square-tiled surfaces have normal genus distribution and cover all integer vectors.
problem Distribution and properties of random square-tiled surfaces.
method Randomizing model and local central limit theorem for genus.
result The distribution of the genus is asymptotically normal and contains all primitive integer vectors.
Randomization is minimax-optimal for variance in experimental design, even with structure.
problem Designing optimal randomized experiments for variance minimization.
method Analyzing permutation symmetric and non-symmetric sets of outcomes, proposing inference-constrained MSOD.
result Randomization is minimax-optimal for variance, even with structure, and requires uniformity constraints for Fisher's exact test.
This paper explains CART random forests using stochastic control theory.
problem Understanding the inner workings of CART random forests.
method Developed a stochastic-control perspective on CART random forests, interpreting feature subsampling as a random feasible action set and the split rule as a policy.
result Established that the CART policy is locally stabilizing but globally suboptimal for the forest objective.
Random forest can be adapted for open-set recognition with improved performance.
problem Handling unknown classes in real-world classification tasks.
method Incorporating distance metric learning and distance-based open-set recognition into random forest.
result The proposed method outperforms state-of-the-art open-set recognition methods.
Simple conditions for comonotonic additive risk measures from acceptance sets.
problem Conditions for comonotonic additive risk measures from acceptance sets.
method Conditions on acceptance sets for induced comonotonic additive risk measures.
result Acceptance sets induce comonotonic additive risk measures if and only if the acceptance sets and their complements are stable under convex combinations of comonotonic random variables.
The paper verifies the robustness of classifier ensembles against randomized attacks.
problem Ensuring classifier ensembles are robust against arbitrary randomized attacks.
method Formal verification procedure using SMT and MILP encodings to assess robustness.
result Proves the NP-hardness of the robustness-checking problem and provides upper bounds on attack sets.
Dimension reduction is the process of embedding high-dimensional data into a lower dimensional space to facilitate its analysis. In the Euclidean setting, one fundamental technique for dimension reduction is to apply a random linear map to the data. This dimension reduction procedure succeeds when it preserves certain …
We develop correlated random measures, random measures where the atom weights can exhibit a flexible pattern of dependence, and use them to develop powerful hierarchical Bayesian nonparametric models. Hierarchical Bayesian nonparametric models are usually built from completely random measures, a Poisson-process based c…
Paper generalizes bipolar theorems for non-negative random variables.
problem Problems with existing bipolar theorems under stronger assumptions.
method Generalizes existing theorems in a robust probabilistic framework.
result Provides necessary and sufficient conditions for bipolar representation.
The main goal of this article is to understand how the length spectrum of a random surface depends on its genus. Here a random surface means a surface obtained by randomly gluing together an even number of triangles carrying a fixed metric. Given suitable restrictions on the genus of the surface, we consider the number…
Random feature matrices' singular values concentrate near their full expectation in high dimensions.
problem Characterizing the spectra of random feature matrices for regression problems.
method Analyzing two settings of input variables (random or well-separated) with conditions on dimension, complexity ratio, and sampling variance.
result The singular values of random feature matrices concentrate near their full expectation and near one with high probability.
Brooks and Makover introduced an approach to random Riemann surfaces based on associating a dense set of them - Belyi surfaces - with random cubic graphs. In this paper, using Bollobas model for random regular graphs, we examine the topological structure of these surfaces, obtaining in particular an estimate for the ex…
New method clusters hypergraphs using weighted random walks and Laplacians.
problem Clustering hypergraph data with edge-dependent weights.
method Random walks with edge-dependent vertex weights, constructing hypergraph Laplacians for clustering.
result Proposed methods outperform existing hypergraph clustering algorithms.
Study reveals a universal formula for knotting in random equilateral polygons.
problem Probability of knotting in equilateral random polygons.
method Extensive Monte Carlo simulations with improved algorithms and knot invariants.
result A universal scaling formula for knotting probability with number of edges, involving exponential and power law factors.
Formula found for probability of random triangles on flat tori being homotopically trivial.
problem Calculating the probability of random triangles on flat tori being homotopically trivial.
method Reduced problem to new invariant of measurable sets in the plane unchanged by area-preserving affine transformations.
result Probability is minimized on rectangular tori and maximized on regular hexagonal tori.
New method robust to random distributional shifts in prediction.
problem Random distributional shifts in real-world settings.
method Hybrid approach combining long-term and proxy outcomes.
result Hybrid approach yields lower mean-squared error than current methods.
Study spectral gaps and bass notes of random hyperbolic 3-orbifolds.
problem Investigate spectral properties of random hyperbolic 3-orbifolds.
method Analyze two models of random hyperbolic 3-orbifolds related to Apollonian and super Apollonian groups.
result Explicit spectral gaps determined for random orbifolds.
Sublinear functionals of random variables are known as sublinear expectations; they are convex homogeneous functionals on infinite-dimensional linear spaces. We extend this concept for set-valued functionals defined on measurable set-valued functions (which form a nonlinear space), equivalently, on random closed sets. …
We introduce sparse random projection, an important dimension-reduction tool from machine learning, for the estimation of discrete-choice models with high-dimensional choice sets. Initially, high-dimensional data are compressed into a lower-dimensional Euclidean space using random projections. Subsequently, estimation …
Alpha-trimming prunes trees in random forests to improve predictive performance.
problem Improving predictive performance of random forests by locally adaptive tree pruning.
method Alpha-trimming is a fast pruning algorithm that prunes trees in a random forest based on signal-to-noise ratio, controlled by a tuning parameter.
result Alpha-trimming often lowers mean squared prediction error compared to fully grown random forests.
Random forests remain among the most popular off-the-shelf supervised machine learning tools with a well-established track record of predictive accuracy in both regression and classification settings. Despite their empirical success as well as a bevy of recent work investigating their statistical properties, a full and…
Introduces RPU to explain randomization preference in dynamic settings.
problem Explains preference for randomization in dynamic investment problems.
method Introduces recursive perturbed utility (RPU) to incorporate randomization preference.
result Proves RPU-optimal portfolio policy is Gaussian and can be expressed in closed form.
The paper studies random dynamical systems of polynomial automorphisms on C^2 and finds mean stability.
problem Random dynamical systems of polynomial automorphisms on C^2.
method Generic random dynamical systems of polynomial automorphisms are shown to have mean stability.
result A generic random dynamical system of polynomial automorphisms on C^2 has mean stability.
Sparse random features improve accuracy in data-scarce settings.
problem Limited accuracy of random feature methods in data-scarce applications.
method Sparse random feature expansion using compressive sensing.
result Improved generalization bounds for sparse random features.
New algorithms find half-optimal independent sets in sparse graphs.
problem Finding large independent sets in sparse random graphs.
method Low-degree polynomial algorithms.
result Low-degree polynomial algorithms can find independent sets of half-optimal size.
Paper introduces PCP for efficient, reliable predictive inference.
problem Developing reliable predictive inference methods for target variables.
method Probabilistic conformal prediction using conditional random samples.
result PCP provides sharper predictive sets compared to existing methods.
New random forest method provides optimal rates and confidence bands.
problem Improving random forest regression rates and constructing confidence bands.
method Proposed Ehrenfest centered purely random forests achieve optimal rates; used Gaussian approximation for supremum of empirical processes.
result Explicit asymptotic uniform confidence bands constructed for both random forest types.
Tian's theorem connects Chern classes of bundles to random section zeros and degeneracy sets.
problem Understanding the distribution of zeros and degeneracy sets of random holomorphic sections.
method Analyzing the pullback of Chern classes and computing currents of integration.
result The limit distribution of zeros of random sections is determined by the Chern form.
Estimates set overlap and similarity using random samples.
problem Estimating set overlap and similarity with limited data.
method Binomial model for predicting set overlap, comparing to previous methods.
result Binomial model provides better estimates with small sample sizes.
Random Transformers behave like polynomial models in ICL with asymptotic growth.
problem Understanding in-context learning capabilities of pretrained Transformers.
method Asymptotic analysis of a random Transformer with a fixed first layer and a trained second layer, considering growth in context length, input dimension, hidden dimension, and training parameters.
result The random Transformer's ICL error is equivalent to a finite-degree Hermite polynomial model.
We study the asymptotic properties of the conormal cycle of nodal sets associated to a random superposition of eigenfunctions of the Laplacian on a smooth compact Riemannian manifold without boundary. In the case where the dimension is odd, we show that the expectation of the corresponding current of integration equidi…
The paper bounds solutions to complex optimization problems with uncertain data.
problem Distributionally robust optimization problems with multivariate uncertainty sets.
method Conditions and bounds derived for multivariate and univariate Wasserstein distances, Bregman-Wasserstein divergences, and signed Choquet integrals.
result Computable lower and upper bounds for DRO problems, derived from scalar-valued aggregation functions and Wasserstein distances.
TOO optimizes stochastic epidemiological models by finding both parameter settings and random seeds.
problem Calibrating stochastic epidemiological models to match empirical observations.
method Gaussian process surrogates and Thompson sampling for optimization.
result Produces actual trajectories consistent with ground truth.
Randomization helps verify if data mining results are due to inherent patterns.
problem Verify if data mining results are due to inherent patterns or coincidental findings.
method Metropolis sampling based on local swaps to randomize data while preserving discovered patterns.
result Randomized data often reveals that clustering results imply frequent pattern discovery.
An algorithm finds optimal covariates for blocking in randomized experiments.
problem Minimizing variance in causal effect estimates from heterogeneous data.
method Using causal graphs, an algorithm identifies optimal covariates for blocking.
result An efficient algorithm reduces variance in causal effect estimates.