A new classification rule for FDA improves classification performance by accounting for unequal covariance matrices.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Hexagonal tilings minimize perimeter with unequal volumes.
Combines BART and Gaussian process for spatial covariate prediction with uncertainty.
New method for MMD with unequal sample sizes improves test power.
PAMA learns covariate importance for better matching in observational studies.
New algorithm clusters Gaussian mixtures with unknown covariance efficiently.
Bayesian approach learns linear networks from high-dimensional data.
A method for rank verification in multivariate Gaussian data, improving on existing approaches.
We consider the problem of high-dimensional classification between the two groups with unequal covariance matrices. Rather than estimating the full quadratic discriminant rule, we propose to perform simultaneous variable selection and linear dimension reduction on original data, with the subsequent application of quadr…
Support vector regression (SVR) has been widely used to reduce the high computational cost of computer simulation. SVR assumes the input parameters have equal sample sizes, but unequal sample sizes are often encountered in engineering practices. To solve this issue, a new prediction approach based on SVR, namely as hig…
This study examines the topology of singularities in optimal semicouplings between unequal spaces.
Clinical data for ambulatory care, which accounts for 90% of the nations healthcare spending, is characterized by relatively small sample sizes of longitudinal data, unequal spacing between visits for each patient, with unequal numbers of data points collected across patients. While deep learning has become state-of-th…
Paper explores Bayes rule for Gaussian mixtures with missing data, outperforming supervised classifiers.
Recent work shows unequal performance of commercial face classification services in the gender classification task across intersectional groups defined by skin type and gender. Accuracy on dark-skinned females is significantly worse than on any other group. In this paper, we conduct several analyses to try to uncover t…
Enhances projection pursuit tree classifier with visual diagnostics for better multi-class classification.
New method uses causal thinking to make AI fairer decisions.
IMPACT optimizes LLM compression by focusing on activation importance, reducing model size up to 55.4%.
The paper analyzes trade dynamics among G7 countries, revealing unequal exchange and degenerate equilibrium states.
Sharp spectral theorems and isoperimetric inequalities for manifolds with nonnegative Ricci curvature.
This paper introduces forward-looking measures of the network connectedness of fears in the financial system, arising due to the good and bad beliefs of market participants about uncertainty that spreads unequally across a network of banks. We argue that this asymmetric network structure extracted from call and put tra…
Robust fuzzy clustering for EEG driver alertness with outlier detection.
Study examines how taxes affect wealth inequality in economic models.
Controversies around race and machine learning have sparked debate among computer scientists over how to design machine learning systems that guarantee fairness. These debates rarely engage with how racial identity is embedded in our social experience, making for sociological and psychological complexity. This complexi…
We consider the problem of closeness testing for two discrete distributions in the practically relevant setting of \emph{unequal} sized samples drawn from each of them. Specifically, given a target error parameter , independent draws from an unknown distribution and draws from an unkno…
This Chapter reviews statistical models for the probability distribution of money developed in the econophysics literature since the late 1990s. In these models, economic transactions are modeled as random transfers of money between the agents in payment for goods and services. Starting from the initially equal distrib…
The covering spectrum is a geometric invariant of a Riemannian manifold, more generally of a metric space, that measures the size of its one-dimensional holes by isolating a portion of the length spectrum. In a previous paper we demonstrated that the covering spectrum is not a spectral invariant of a manifold in dimens…
A new framework for robust transfer learning that avoids negative transfer in domains with unequal information.
Bayesian model tackles spatial count data issues with flexible non-parametric techniques.
We study the misclassification error for community detection in general heterogeneous stochastic block models (SBM) with noisy or partial label information. We establish a connection between the misclassification rate and the notion of minimum energy on the local neighborhood of the SBM. We develop an optimally weighte…
Research into time series classification has tended to focus on the case of series of uniform length. However, it is common for real-world time series data to have unequal lengths. Differing time series lengths may arise from a number of fundamentally different mechanisms. In this work, we identify and evaluate two cla…
A statistical generalization is made of microeconomics in the spirit of going from classical to statistical mechanics. The price and quantity of every commodity1 traded in the market, at each instant of time, is considered to be an independent random variable: all prices and quantities are considered to be stochastic p…
In 2016, the majority of full-time employed women in the U.S. earned significantly less than comparable men. The extent to which women were affected by gender inequality in earnings, however, depended greatly on socio-economic characteristics, such as marital status or educational attainment. In this paper, we analyzed…
Mathai-Quillen forms are used to give an integral formula for the Lefschetz number of a smooth map of a closed manifold. Applied to the identity map, this formula reduces to the Chern-Gauss-Bonnet theorem. The formula is computed explicitly for constant curvature metrics. There is in fact a one-parameter family of inte…
SGD with constant stepsize converges to a non-Gaussian limit near flat minima.
Developing tools for computing string amplitudes with hyperbolic vertices.
AF improves classification models by adaptively weighting trees.
We present a technique for clustering categorical data by generating many dissimilarity matrices and averaging over them. We begin by demonstrating our technique on low dimensional categorical data and comparing it to several other techniques that have been proposed. Then we give conditions under which our method shoul…
New findings on maximizing noise stability in partitions of Gaussian space.
We study the effect of the social stratification on the wealth distribution on a system of interacting economic agents that are constrained to interact only within their own economic class. The economical mobility of the agents is related to its success in exchange transactions. Different wealth distributions are obtai…
We propose new positive definite kernels for permutations. First we introduce a weighted version of the Kendall kernel, which allows to weight unequally the contributions of different item pairs in the permutations depending on their ranks. Like the Kendall kernel, we show that the weighted version is invariant to rela…
The outcome of a functional genomics pipeline is usually a partial list of genomic features, ranked by their relevance in modelling biological phenotype in terms of a classification or regression model. Due to resampling protocols or just within a meta-analysis comparison, instead of one list it is often the case that …
A widely applied diversification paradigm is the naive diversification choice heuristic. It stipulates that an economic agent allocates equal decision weights to given choice alternatives independent of their individual characteristics. This article provides mathematically and economically sound choice theoretic founda…
The intermarket analysis, in particular the lead-lag relationship, plays an important role within financial markets. Therefore a mathematical approach to be able to find interrelations between the price development of two different financial underlyings is developed in this paper. Computing the differences of the relat…
Importance sampling is often used in machine learning when training and testing data come from different distributions. In this paper we propose a new variant of importance sampling that can reduce the variance of importance sampling-based estimates by orders of magnitude when the supports of the training and testing d…
Rank aggregation systems collect ordinal preferences from individuals to produce a global ranking that represents the social preference. Rank-breaking is a common practice to reduce the computational complexity of learning the global ranking. The individual preferences are broken into pairwise comparisons and applied t…
In certain situations that shall be undoubtedly more and more common in the Big Data era, the datasets available are so massive that computing statistics over the full sample is hardly feasible, if not unfeasible. A natural approach in this context consists in using survey schemes and substituting the "full data" stati…
Study uniform rates for estimating Gaussian mixtures without separation assumption.
We consider a Bayesian framework for estimating a high-dimensional sparse precision matrix, in which adaptive shrinkage and sparsity are induced by a mixture of Laplace priors. Besides discussing our formulation from the Bayesian standpoint, we investigate the MAP (maximum a posteriori) estimator from a penalized likel…