New method trains complex models on impure data, matching simulation performance.
problem Training complex models on impure data with limited labeled samples.
method Weak supervision techniques for high-dimensional data.
result Complex high-dimensional classifiers trained on impure mixtures perform comparably to pure samples.
The study examines how altering impurity functions influences optimal splits in binary classification trees.
problem Understanding how altering impurity functions affects optimal splits in binary classification trees.
method Investigates how skewing impurity functions biases optimal splits towards isolating points of a particular class.
result A necessary and sufficient condition for skewing an impurity function to bias optimal splits towards isolating points of a particular class is provided.
Bregman perspective on CART provides a unified framework for impurity measures.
problem Unifying impurity measures in CART
method Bregman divergence approach
result Unified framework for impurity measures
New method classifies spinor orbits in dimensions up to 14.
problem Classify orbits of the spin group in semi-spinor spaces.
method Use pure spinors and combinatorial constraints to classify orbits.
result Classification of spinor orbits in dimensions up to 14.
We consider supersymmetric gauge theories with impurities in various dimensions. These systems arise in the study of intersecting branes. Unlike conventional gauge theories, the Higgs branch of an impurity theory can have compact directions. For models with eight supercharges, the Higgs branch is a hyperKahler manifold…
New method classifies mixtures without labels, recovering latent classes.
problem Classification without reliable instance-level labels.
method Posterior simplex geometry for multiclass learning.
result Classifier trained on mixture identities recovers latent classes and their proportions.
A new method estimates uncertainty without explicit prediction models.
problem Costly data acquisition in machine learning.
method Distance-weighted Class Impurity method for uncertainty estimation.
result Distance-weighted Class Impurity effectively estimates uncertainty without prediction models.
A hybrid impurity measure balances theoretical soundness and computational efficiency.
problem Developing a robust impurity measure for decision trees.
method Integrates Tsallis entropy with an exponential polarization component.
result Simple parametric measures outperform ITC, but ITC variants are competitive with strong theoretical guarantees.
The paper studies bias and adaptivity of CART regression trees.
problem Bias and adaptivity of CART regression trees.
method Derives an interesting connection between bias and MDI measure of variable importance.
result Decision trees with CART have small bias and are adaptive to signal strength and direction.
This paper analyzes Mean Decrease Impurity (MDI) variable importance in random forests.
problem Lack of interpretability in random forest variable importances.
method Analysis of Mean Decrease Impurity (MDI) in random forests.
result MDI provides a variance decomposition of the output when variables are independent and there are no interactions.
Machine learning methods are applied to finding the Green's function of the Anderson impurity model, a basic model system of quantum many-body condensed-matter physics. Different methods of parametrizing the Green's function are investigated; a representation in terms of Legendre polynomials is found to be superior due…
The paper analyzes the convergence of CART under a SID condition, improving previous results.
problem Investigating the convergence rate of CART under a sufficient impurity decrease condition.
method Established an upper bound on prediction error under SID condition, introduced easily verifiable conditions.
result Improved convergence rate of CART under SID condition, demonstrated examples of error bound limitations.
AquaSight detects water impurity using deep learning.
problem Water pollution and lack of affordable water quality assessment.
method Convolutional Neural Networks for automated water impurity detection.
result Deep learning model achieved 96% accuracy in detecting water contamination.
Recommender systems improve quantum Monte Carlo simulations.
problem Efficiency of quantum Monte Carlo methods without sacrificing accuracy.
method Quantum to classical mapping and molecular simulation techniques.
result Classical molecular gas model reproduces quantum distributions efficiently.
Optimizes harvesting in biopharmaceutical fermentation with limited data.
problem High variability in fermentation outputs due to model risk with small data.
method Stochastic model, Bayesian approach, Markov decision process.
result Improves fermentation output and reduces variability.
Global stability bounds for matrix frames in phase retrieval problems.
problem Phase retrieval for matrix frames in various applications.
method Computable global stability bounds for the quasi-linear analysis map β, using Whitney stratification of positive semidefinite matrices of low rank.
result Novel conditions for a frame to be generalized phase retrievable.
A new metric assesses spatio-temporal forecast quality using Gini regularized Optimal Transport.
problem Evaluating the quality of spatio-temporal forecasts.
method Gini-regularized Optimal Transport (OT) problem, using Gini impurity function as a regularizer.
result The Gini regularized OT problem converges to the classical OT problem, offering a numerically more stable algorithm.
Novel loss functions improve decision tree learning from noisy data.
problem Training decision trees with noisy labels.
method Introducing distribution losses and a new negative exponential loss.
result The negative exponential loss leads to efficient and robust decision tree learning.
New local MDI variable importances derived from global scores match Shapley values.
problem Local feature relevance in tree-based models.
method Deriving local MDI importance measure from global scores and linking it to Shapley values.
result Local MDI importances have a natural connection with Shapley values.
We present ADHM-Nahm data for instantons on the Taub-NUT space and encode these data in terms of Bow Diagrams. We study the moduli spaces of the instantons and present these spaces as finite hyperkahler quotients. As an example, we find an explicit expression for the metric on the moduli space of one SU(2) instanton. W…
This paper optimizes high-dimensional oblique splits for decision trees, enhancing performance and computational efficiency.
problem Enhancing decision tree performance and computational efficiency in high-dimensional data.
method Established Sufficient Impurity Decrease (SID) convergence for s0-sparse oblique splits, proposing progressive trees for iterative refinement. result Demonstrated that SID function class expands with s0-sparsity, enabling capture of complex data-generating processes. Self-poisoning in adaptive OOD detectors is explained with a sharp threshold theory and certified calibration.
problem Self-poisoning in adaptive OOD detectors.
method Modeling bank impurity as a generalized Pólya urn, proving almost-sure convergence to a mean-field equilibrium.
result A certified admission gate removes the transition at every contamination rate, controlling false positives label-free.
Paper explores how unsupervised learning can be understood through linear algebra concepts.
problem Understanding unsupervised learning through linear algebra concepts.
method Introducing the concept of linearly independent populations and using them to solve for prevalence values.
result Unsupervised learning can be realized as a generalization of supervised learning.
Study uses random forest to detect unlawful insider trading in financial data.
problem Detecting and identifying unlawful insider trading in complex financial data.
method Integrates PCA-RF and standalone RF models with semi-manually labeled transactions.
result 96.43% accurate classification of transactions, 95.47% lawful, 98.00% unlawful.
Novel approach for creating interpretable classifiers using bilevel optimization of split-rules in NLDTs.
problem Creating highly accurate and easily interpretable classifiers for practical applications.
method Representing classifiers as assemblies of simple mathematical rules using NLDTs with evolutionary bilevel optimization.
result The approach ensures interpretability while achieving high accuracy on various classification problems.
Develops RF-GLS for binary geospatial data.
problem Challenges in extending RF to binary geospatial data.
method Proposes RF-GLS for binary data, embedding it in generalized mixed effects models.
result Establishes consistency of RF-GP for mean function and covariate effect estimation.
New method compresses Green's function data efficiently.
problem Efficiently representing complex correlation functions.
method Intermediate representation (IR) of analytical continuation.
result IR yields significantly compact form of correlation functions.
We seek decision rules for prediction-time cost reduction, where complete data is available for training, but during prediction-time, each feature can only be acquired for an additional cost. We propose a novel random forest algorithm to minimize prediction error for a user-specified {\it average} feature acquisition b…
Regularizes decision trees to reduce inference time by up to 4x with minimal accuracy loss.
problem Optimizing decision tree execution time on resource-constrained devices.
method Regularizes impurity computation during CART algorithm training to favor highly asymmetric distributions.
result Reduces inference time by up to 4x with minimal accuracy loss.
Proposes MRF for consistency and privacy in RF.
problem Insufficient theoretical understanding of RF's consistency and privacy.
method Introduces MRF with multinomial distributions for feature and value selection.
result Proves MRF's consistency and analyzes its privacy within differential privacy.
Deep neural network speeds up plasma tomography by several orders of magnitude.
problem Slow computation of plasma tomographic reconstructions.
method Trained a deep neural network on a large dataset of tomograms.
result Deep neural network can compute all reconstructions in seconds.
We study both the continuous model and the discrete model of the integer quantum Hall effect on the hyperbolic plane in the presence of disorder, extending the results of an earlier paper [CHMM]. Here we model impurities, that is we consider the effect of a random or almost periodic potential as opposed to just periodi…
Study calculates tail risk for various mixture distributions.
problem Estimating tail risk for complex distribution mixtures.
method Analyzes tail conditional expectation for location-scale mixtures of elliptical distributions.
result Developed methods for calculating tail risk in various distributions.
Enhances mixture models with classifier-defined weights.
problem Density evaluation and sampling in mixture models.
method Introduces Classifier Weighted Mixtures (CWM) with functional weights.
result Improves expressivity in variational estimation without increasing complexity.
Two approaches improve parameter learning in various mixture models.
problem Parameter learning in mixture models.
method Complex-analytic and algebraic-combinatorial methods.
result Improved sample sufficiency for parameter estimation in specific mixture models.
We consider unsupervised estimation of mixtures of discrete graphical models, where the class variable corresponding to the mixture components is hidden and each mixture component over the observed variables can have a potentially different Markov graph structure and parameters. We propose a novel approach for estimati…
New algorithm reduces discrimination in predictions.
problem Tackles potential discrimination in AI predictions.
method Integrates fairness adjustments into tree-building process.
result Reduces discriminatory predictions without significant loss in accuracy.
Paper introduces Normalized Wasserstein measure for better handling of imbalanced mixture distributions.
problem Wasserstein distance fails for mixture distributions with imbalanced proportions.
method Introduce mixture proportions as optimization variables to normalize Wasserstein formulation.
result Normalized Wasserstein measure leads to significant performance gains for mixture distributions.
Improved VB algorithm for NIG mixtures outperforms Gaussian mixtures for non-Gaussian data.
problem Clustering non-Gaussian data, especially heavy-tailed and asymmetric.
method Proposed an improved VB algorithm for NIG mixture models and extended Dirichlet process mixture models.
result Outperforms Gaussian mixtures and existing NIG mixture models, especially for highly non-normative data.
When estimating finite mixture models, it is common to make assumptions on the mixture components, such as parametric assumptions. In this work, we make no distributional assumptions on the mixture components and instead assume that observations from the mixture model are grouped, such that observations in the same gro…
The two most extended density-based approaches to clustering are surely mixture model clustering and modal clustering. In the mixture model approach, the density is represented as a mixture and clusters are associated to the different mixture components. In modal clustering, clusters are understood as regions of high d…
Optimal mixtures of generative models outperform individual models on image datasets.
problem Selecting the best single model from a group of trained generative models.
method Formulated a quadratic optimization problem and proposed the Mixture-UCB algorithm for efficient selection.
result Mixture of generative models achieves better evaluation scores than individual models on benchmark datasets.
New bounds on sample size for identifying mixture models with grouped samples.
problem Identifying mixture models with minimal sample size.
method Generalized identifiability bounds for mixture models with grouped samples.
result Identifiability with (2m−1)/(k−1) samples per group, with no improvement possible. Bayesian approach learns nonparametric mixture components from heterogeneous data.
problem Realistic modeling of heterogeneous data populations with nonparametric mixture components.
method Bayesian nonparametric modeling using Dirichlet process mixture priors.
result Posterior contraction rates for component densities are nearly polynomial, improving over deconvolution methods.
DGMM uses deep layers of Gaussian mixtures for flexible data modeling.
problem Efficiently modeling complex data relationships.
method Deep Gaussian Mixture Models (DGMM) with nested mixtures of linear models and factor models.
result DGMM provides a flexible nonlinear model for data description.
Robust learning mixtures of linear regressions improve robustness.
problem Improving robustness in learning mixtures of linear regressions.
method Connecting mixtures of linear regressions and mixtures of Gaussians with thresholding for a quasi-polynomial time algorithm.
result The algorithm has significantly better robustness than previous results.
Spatially constrained Gaussian mixture models reduce covariance complexity.
problem High dimensionality in finite mixture models for spatial data.
method Spatial covariance constraint with only four free parameters.
result Improves clustering of multi-way spatial data and inference of spatial patterns.
This paper compares EM and GD in two-component mixture models, finding EM escapes bad local optima more reliably.
problem Understanding the convergence of EM and GD in mixture models, especially in regions where one component is missing.
method Analyzing regions called one-cluster regions in two-component mixture models of Gaussians and Bernoullis, comparing the propensity of EM and GD to converge to these regions.
result EM escapes one-cluster regions exponentially fast, while GD escapes them linearly fast, indicating EM is less likely to converge to bad local optima.