Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

2356 · Feb 202619922001200920182026
48 results for Gini impurity

A new metric assesses spatio-temporal forecast quality using Gini regularized Optimal Transport.

problem Evaluating the quality of spatio-temporal forecasts.
method Gini-regularized Optimal Transport (OT) problem, using Gini impurity function as a regularizer.
result The Gini regularized OT problem converges to the classical OT problem, offering a numerically more stable algorithm.

The study examines how altering impurity functions influences optimal splits in binary classification trees.

problem Understanding how altering impurity functions affects optimal splits in binary classification trees.
method Investigates how skewing impurity functions biases optimal splits towards isolating points of a particular class.
result A necessary and sufficient condition for skewing an impurity function to bias optimal splits towards isolating points of a particular class is provided.

Study uses random forest to detect unlawful insider trading in financial data.

problem Detecting and identifying unlawful insider trading in complex financial data.
method Integrates PCA-RF and standalone RF models with semi-manually labeled transactions.
result 96.43% accurate classification of transactions, 95.47% lawful, 98.00% unlawful.

We consider supersymmetric gauge theories with impurities in various dimensions. These systems arise in the study of intersecting branes. Unlike conventional gauge theories, the Higgs branch of an impurity theory can have compact directions. For models with eight supercharges, the Higgs branch is a hyperKahler manifold…

1998-04-03abs ↗pdf ↗

New method trains complex models on impure data, matching simulation performance.

problem Training complex models on impure data with limited labeled samples.
method Weak supervision techniques for high-dimensional data.
result Complex high-dimensional classifiers trained on impure mixtures perform comparably to pure samples.

Paper proposes Gini distance statistics for estimating feature-label dependence.

problem Identifying statistical dependence between features and categorical labels.
method Generalized Gini distance in RKHS for feature-label dependence estimation.
result Gini distance statistics converge faster and have tighter error bounds than distance covariance.

The paper examines stochastic ordering of Gini indexes for multivariate elliptical risks.

problem Stochastic ordering of Gini indexes for multivariate elliptical risks.
method Established conditions for monotonicity of Gini index in usual stochastic order.
result Suitable conditions for multivariate elliptical risks generalize those for multivariate normal risks.

The paper examines the stability of binary choice models using Gini index and scoring indicators.

problem Stability and discriminatory power of binary choice models.
method Derives the real Gini index and incorporates PSI and KS statistics into the model.
result The real Gini index should be less than the calculated Gini index when the population distribution changes.

Improved decision tree learning guarantees for complex functions.

problem Achieving provable guarantees for decision tree induction with complex target functions.
method Introduces a new splitting criterion that considers correlations between target function and subsets of attributes.
result Proves provable guarantees for all target functions with respect to the uniform distribution, circumventing previous impossibility results.

Paper proves convergence of Gini index to equilibrium in Wasserstein distance.

problem Proving convergence of Gini index to equilibrium in Wasserstein distance.
method Analyzes Gini index as Lyapunov functional and proves convergence in Wasserstein distance.
result Proves convergence of Gini index to equilibrium in Wasserstein distance.

Kernel methods are popular in clustering due to their generality and discriminating power. However, we show that many kernel clustering criteria have density biases theoretically explaining some practically significant artifacts empirically observed in the past. For example, we provide conditions and formally prove the…

2017-05-16abs ↗pdf ↗

A hybrid impurity measure balances theoretical soundness and computational efficiency.

problem Developing a robust impurity measure for decision trees.
method Integrates Tsallis entropy with an exponential polarization component.
result Simple parametric measures outperform ITC, but ITC variants are competitive with strong theoretical guarantees.

DAC improves associative classification for very large datasets with high scalability and quality.

problem Handling large datasets with many categorical features.
method DAC uses ensemble learning, parallel processing, and pruning techniques.
result DAC outperforms state-of-the-art solutions in prediction quality and execution time.

Analytical results bound the approach to oligarchy in a modified asset exchange model.

problem Analyzing economic inequality in a modified asset exchange model.
method Analytical results using Gini coefficient and differential inequalities.
result The Gini coefficient is bounded by a first-order differential inequality.

Study on Gini estimation for fat-tailed data, showing bias and proposing corrections.

problem Estimating Gini index under infinite variance data.
method Analysis of nonparametric and maximum likelihood estimators, focusing on phase transitions and tail index effects.
result Maximum likelihood estimation outperforms nonparametric methods for fat-tailed data.

The paper analyzes worst-case distortion risk metrics and weighted entropy under partial information.

problem Analyzing worst-case distortion risk metrics and weighted entropy with limited information.
method General distributions, partial information (mean and variance), various entropies and risk measures.
result Provides worst-case results for distortion risk metrics and weighted entropy.

Direct measurements of Gini coefficients by conventional arithmetic calculations are a poor estimator, even if paradoxically, they include the entire population, as because of super-additivity they cannot lend themselves to comparisons between units of different size, and intertemporal analyses are vitiated by the popu…

2015-10-16abs ↗pdf ↗

The paper studies bias and adaptivity of CART regression trees.

problem Bias and adaptivity of CART regression trees.
method Derives an interesting connection between bias and MDI measure of variable importance.
result Decision trees with CART have small bias and are adaptive to signal strength and direction.

This paper analyzes Mean Decrease Impurity (MDI) variable importance in random forests.

problem Lack of interpretability in random forest variable importances.
method Analysis of Mean Decrease Impurity (MDI) in random forests.
result MDI provides a variance decomposition of the output when variables are independent and there are no interactions.

The paper analyzes the convergence of CART under a SID condition, improving previous results.

problem Investigating the convergence rate of CART under a sufficient impurity decrease condition.
method Established an upper bound on prediction error under SID condition, introduced easily verifiable conditions.
result Improved convergence rate of CART under SID condition, demonstrated examples of error bound limitations.

Two regularization techniques improve GCNN explainability and preference from chemists.

problem Difficulty in rationalizing molecular graph neural network predictions.
method Batch Representation Orthonormalization (BRO) and Gini regularization applied during GCNN training.
result Regularization improves GCNN attribution methods and preference from chemists.

The standard deviation and Gini mean difference order based on tail behavior.

problem Ordering between standard deviation and Gini mean difference for real-valued risks.
method Analysis of the mean excess function of the pairwise difference XX|X - X'|.
result Dominance regimes of SD and GMD are determined by tail behavior of the distribution.

This paper fills in local bounds for Spearman's footrule and Gini's gamma measures of association.

problem Local bounds for bivariate copulas with respect to Spearman's footrule and Gini's gamma measures.
method Computing quasi-copulas that are not copulas for certain values of the measures.
result Presented local bounds for Spearman's footrule and Gini's gamma measures.

Wealth redistribution through Fokker-Planck equation controls preserves Gini coefficient.

problem Preserving Gini coefficient through proportional wealth tax.
method Formulating optimal redistribution as a control problem for Fokker-Planck equation.
result Progressive taxes redistribute within policy-relevant timescales.

Socio-economic inequality is characterized from data using various indices. The Gini (gg) index, giving the overall inequality is the most common one, while the recently introduced Kolkata (kk) index gives a measure of 1k1-k fraction of population who possess top kk fraction of wealth in the society. Here, we show t…

2016-06-10abs ↗pdf ↗

The paper calculates bounds for risk metrics and entropies under partial information constraints.

problem Analyzing risk metrics and entropies for unimodal, symmetric distributions with limited information.
method Develops lower and upper bounds for worst-case distortion riskmetrics and weighted entropy for unimodal, symmetric distributions with known mean and variance.
result Sharp upper bounds for distortion riskmetrics and weighted entropy for symmetric distributions.

The article compares predictor importance in classification problems with categorical outcomes.

problem Comparing predictor importance in classification problems with categorical response variables.
method The approach is based on the categorical Gini correlation (CGC) and tests differences in CGCs across predictor groups.
result The proposed methodology accommodates predictors of arbitrary and unequal dimensions and allows for dependence between predictor groups.