A method to simplify complex high-dimensional data visualization.
problem Difficult interpretation of linear projections in high-dimensional data.
method Decomposition of linear projections into axis-aligned projections using Dempster-Shafer theory.
result Linear projections can be effectively represented by a sparse set of axis-aligned projections, revealing more intuitive insights.
Decision forests, including Random Forests and Gradient Boosting Trees, have recently demonstrated state-of-the-art performance in a variety of machine learning settings. Decision forests are typically ensembles of axis-aligned decision trees; that is, trees that split only along feature dimensions. In contrast, many r…
New method for sparse PCA using random projections, non-iterative and fast.
problem Sparse principal component analysis (PCA)
method Axis-aligned random projections of sample covariance matrix
result Non-iterative method achieves optimal convergence rate in polynomial time
Paper develops a method to identify feature subspaces contributing to local data complexity.
problem Identifying feature subspaces that contribute to local data complexity.
method Develops an estimator of Local Intrinsic Dimension (LID) along axis projections to identify feature subspaces.
result Preliminary evidence suggests LID decomposition can indicate axis-aligned data subspaces supporting cluster formation.
Sharp-SSL uses random projections to identify important variables for semi-supervised learning.
problem High-dimensional semi-supervised learning problems.
method Careful aggregation of low-dimensional results from many axis-aligned random projections.
result Sharp-SSL algorithm can recover signal coordinates with high probability.
A new method for unsupervised disentanglement using axis-aligned cliffs.
problem Unsupervised disentanglement of latent factors under nonlinear maps.
method Encouraging axis-aligned discontinuities (cliffs) in the estimated density of factors.
result Cliff method outperforms baselines on disentanglement benchmarks.
We introduce canonical correlation forests (CCFs), a new decision tree ensemble method for classification and regression. Individual canonical correlation trees are binary decision trees with hyperplane splits based on local canonical correlation coefficients calculated during training. Unlike axis-aligned alternatives…
GTBO uses group testing to optimize high-dimensional functions efficiently.
problem Optimizing expensive, high-dimensional functions with limited data.
method Group testing to identify active dimensions, then guide optimization.
result GTBO outperforms state-of-the-art methods on high-dimensional benchmarks.
Random Tessellation Process improves multi-dimensional data analysis.
problem Axis-aligned cuts limit flexibility in space partitioning methods.
method Proposes Random Tessellation Process (RTP) for non-axis aligned cuts.
result Improved accuracies in gene expression data analysis.
Oblique BART improves tree-based predictions.
problem Axis-aligned decision rules in BART can be suboptimal.
method Developed an oblique version of BART using data-adaptive hyperplane partitions.
result Oblique BART outperformed axis-aligned BART and other tree methods on benchmarks.
BO method identifies sparse subspaces for efficient high-dimensional optimization.
problem Efficient optimization of high-dimensional black-box functions.
method Sparse Gaussian process surrogate models on axis-aligned subspaces with Hamiltonian Monte Carlo inference.
result SAASBO achieves excellent performance on synthetic and real-world problems.
New kernel interprets 3D anisotropic data with rotations and improved predictions.
problem Capturing rotated anisotropy in 3D spatial fields.
method Introduces a Lie-algebraic kernel with three principal length-scales and an explicit rotation.
result Posterior recovers rotated anisotropy and improves prediction over axis-aligned kernels.
New technique crafts imperceivable sparse adversarial attacks.
problem Vulnerability of neural networks to adversarial attacks.
method Proposes a black-box technique to minimize l0-distance, integrating componentwise constraints. result Adversarial examples are almost imperceivable and non-detectable.
Tree regularization makes deep models interpretable by approximating them with simple decision trees.
problem Lack of interpretability in deep neural networks.
method Tree regularization to train deep models to resemble compact, axis-aligned decision trees.
result Tree regularized models are easier for humans to interpret without sacrificing accuracy.
New method learns representations for decision forests using input perturbation.
problem Decision forests struggle with raw structured data and lack effective representations.
method Approximate decision forest gradients through input perturbation.
result Effective representation learning for decision forests without structural changes.
GTBO uses group testing to optimize high-dimensional functions efficiently.
problem Challenges in optimizing high-dimensional, expensive functions due to the curse of dimensionality.
method GTBO combines testing and optimization phases to identify active variables and guide efficient optimization.
result GTBO outperforms state-of-the-art methods on high-dimensional optimization tasks.
Bayesian nonparametric method partitions shapes using curves.
problem Capturing complex shapes in multi-dimensional data.
method Proposes a novel spline partitioning approach using curves.
result Demonstrates improved shape modeling compared to existing methods.
New random forest variants achieve optimal performance in high dimensions.
problem Handling dependencies between features in high-dimensional data.
method Using oblique splits in random forests with general split directions.
result Achieved minimax optimal convergence rates in arbitrary dimension.
This work uses stochastic geometry to improve STIT processes in machine learning.
problem Improving STIT processes for efficient and consistent machine learning applications.
method Utilizing tools from stochastic geometry to characterize kernels and obtain consistency results.
result Generalization of STIT processes and their kernels, leading to improved machine learning methods.
Lean 4 formalizes Stokes' theorem for smooth singular cubes.
problem Formalizing Stokes' theorem for singular cubes in arbitrary dimensions.
method Using true differential-form pullback via Frechet derivative, bridging to mathlib4's extDeriv.
result d^2=0 for singular cubical chains, chain-level Stokes extended.
Decision trees and shallow neural networks have different geometric complexities, impacting their interpretability and accuracy.
problem The geometric simplicity of decision boundaries in decision trees conflicts with the approximation capabilities of shallow neural networks.
method Analysis of the Radon total variation (RTV) seminorm to compare geometric complexity of decision regions and neural network approximations.
result Smooth barrier scores can approximate decision regions with finite RTV, but their performance depends on the tube-mass condition near the decision boundary.
Decision Machines embeds decision trees into vector spaces for improved optimization.
problem Overfitting and difficulty in finding optimal decision tree structure.
method Embedding Boolean tests into a binary vector space and representing tree structure as matrices.
result Optimized decision trees with enhanced predictive power.
We introduce a new and improved characterization of the label complexity of disagreement-based active learning, in which the leading quantity is the version space compression set size. This quantity is defined as the size of the smallest subset of the training data that induces the same version space. We show various a…
The paper develops a new method to test if two multidimensional distributions are equivalent or significantly different.
problem Testing equivalence of multidimensional distributions with sub-linear sample complexity.
method Uses generalized A_k distance and Ramsey theory to develop a computationally efficient closeness tester.
result First sub-linear sample complexity closeness tester for multidimensional distributions.
Online BSP-Forest improves space partitioning for large-scale classification and regression.
problem Efficient space partitioning for large-scale classification and regression problems.
method Developed an online BSP-Forest framework that expands space coverage and refines partition structure in real-time.
result Guaranteed universal consistency for both classification and regression problems.
Mixture models are a fundamental tool in applied statistics and machine learning for treating data taken from multiple subpopulations. The current practice for estimating the parameters of such models relies on local search heuristics (e.g., the EM algorithm) which are prone to failure, and existing consistent methods …
Regularized LAEs learn principal components efficiently.
problem Learning optimal linear representations with LAEs.
method Proper regularization schemes (non-uniform ℓ2 and nested dropout).
result Convergence to optimal representation is slow due to ill-conditioning.
New TVD estimator adapts to piecewise constant functions, improving performance.
problem Improving TVD estimator performance for piecewise constant functions.
method Investigates adaptivity of TVD estimator to piecewise constant functions and proposes a data-driven tuning parameter.
result The ideally tuned TVD estimator performs better than in the worst case for piecewise constant functions.
Proposes a new BSP-Tree process for flexible space partition modeling.
problem Limited modelling flexibility of axis-aligned partitions in Mondrian process.
method Introduces a self-consistent Binary Space Partitioning (BSP)-Tree process with oblique cuts.
result Clear inferential improvements over standard Mondrian process and related methods.
New BO method efficiently optimizes high-dimensional functions by automatically selecting variables.
problem Efficiently optimizing functions with high-dimensional domains.
method Exploits variable selection to automatically learn sub-spaces without pre-specified dimensions.
result Empirically validated on synthetic and real problems, demonstrating efficiency.
Boundary effects inflate variance in Gaussian processes, leading to acquisition bias.
problem Boundary-induced acquisition bias in Gaussian processes.
method Traced root cause to geometric mechanism of kernel truncation at domain boundaries.
result Boundary effects create distortion that worsens with dimensionality, affecting acquisition behavior.
Proposes regional tree regularization for interpretable deep models.
problem Lack of interpretability in deep neural networks.
method Encourages deep models to be well-approximated by separate decision trees for predefined regions of the input space.
result Regional tree regularization delivers more accurate predictions than training separate decision trees for each region, while producing simpler explanations.
LassoFlexNet improves deep learning performance on tabular data.
problem Deep learning underperforms tree-based models on tabular data.
method Incorporates five inductive biases and uses Tied Group Lasso for variable selection.
result LassoFlexNet matches or outperforms leading tree-based models on 52 datasets.
Paper defines a new dimension to measure self-directed learning complexity.
problem Understanding self-directed learning complexity in online learning theory.
method Developed a dimension SDdim to characterize self-directed learning mistake-bound. result Calculated SDdim for various concept classes and demonstrated learnability gaps. We compare the sample complexity of private learning [Kasiviswanathan et al. 2008] and sanitization~[Blum et al. 2008] under pure ε-differential privacy [Dwork et al. TCC 2006] and approximate (ε,δ)-differential privacy [Dwork et al. Eurocrypt 2006]. We show that the sample complexity of these tasks under approxima…
This work explores feature learning tradeoffs in neural networks.
problem Resource tradeoffs in neural feature learning.
method Theoretical and experimental investigation of offline sparse parity learning.
result Width improves sample efficiency in sparse feature learning.
Study reduces memory needs for active learning with enriched queries.
problem Expensive labeling costs in active learning.
method Introduces bounded memory active learning through enriched queries, introduces lossless sample compression.
result Can learn classifiers with bounded memory and query optimality.
Optimized coordinate system improves sparse grid regression performance.
problem Sparse grid methods struggle with skewed and rotated coordinates.
method Proposes an optimized coordinate system to reduce effective dimensionality.
result Adaptive sparse grid least squares algorithm benefits from preprocessing.
New method certifies neural network function space norms from point evaluations.
problem Certifying neural network function space norms from point evaluations alone.
method Combining interval arithmetic enclosures, adaptive marking/refinement, and quadrature-based aggregation.
result Certified computation of Lp, W1,p, and W2,p norms. Paper introduces conformal prediction for reliable uncertainty quantification in landmark localization.
problem Systematic underestimation of total predictive uncertainty in landmark localization.
method Conformal prediction framework for multi-output regression, generating flexible prediction regions.
result Methods outperform existing approaches in validity and efficiency across 2D and 3D datasets.
A benchmarking framework for studying data geometry.
problem Generalization and approximation error bounds in deep learning.
method Repurposing and extending dSprites and COIL-20 with additional transformation dimensions and dense, axis-aligned sampling.
result Near-ground-truth accuracy in curvature, reach, and volume estimation.
This paper improves coreset construction for kernel density estimates.
problem Approximating large kernel density estimates with smaller ones.
method Developed a coreset construction algorithm called kernel herding.
result Approximates kernel density estimates with much smaller point sets.
Self-directed learners can minimize mistakes in online classification.
problem Minimizing mistakes in online classification with adaptive prediction order.
method Designing efficient self-directed learners for linear classification.
result Strong separation between worst-order and random-order learning for linear classification.
Low rank tensor decompositions are a powerful tool for learning generative models, and uniqueness results give them a significant advantage over matrix decomposition methods. However, tensors pose significant algorithmic challenges and tensors analogs of much of the matrix algebra toolkit are unlikely to exist because …
Decision forests learn to model text by evaluating categorical-set conditions.
problem Decision forests cannot directly model text features.
method Defined and learned conditions for categorical-set features, enabling efficient text modeling.
result Decision forests can now directly model text features.
The paper defines projective structures for Lie bialgebras and Poisson-Lie groups.
problem Defining projective analogues of Lie bialgebras and Poisson-Lie groups.
method Introducing projective tensor products and adapting classical notions to these structures.
result Every quasi-triangular projective r-matrix gives rise to a projective Banach Lie bialgebra.
Estimates piecewise polynomials and bounded variation functions using optimal decision trees.
problem Estimating piecewise smooth functions in general dimensions.
method Dyadic CART and Optimal Regression Tree (ORT) estimators for piecewise polynomials and bounded variation functions.
result Oracle inequalities and risk bounds for ORT estimators, demonstrating adaptivity and optimality.
Characterizes minor-minimal separating projective planar graphs and their generalizations.
problem Understanding projective planar graphs and their properties.
method Analyzing minors, embeddings, and specific link types.
result Partial characterization of minor-minimal separating projective planar graphs and their generalizations.