New method for MMD with unequal sample sizes improves test power.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Support vector regression (SVR) has been widely used to reduce the high computational cost of computer simulation. SVR assumes the input parameters have equal sample sizes, but unequal sample sizes are often encountered in engineering practices. To solve this issue, a new prediction approach based on SVR, namely as hig…
Clinical data for ambulatory care, which accounts for 90% of the nations healthcare spending, is characterized by relatively small sample sizes of longitudinal data, unequal spacing between visits for each patient, with unequal numbers of data points collected across patients. While deep learning has become state-of-th…
We consider the problem of closeness testing for two discrete distributions in the practically relevant setting of \emph{unequal} sized samples drawn from each of them. Specifically, given a target error parameter , independent draws from an unknown distribution and draws from an unkno…
The paper improves A/B testing for non-Gaussian data, ensuring reliable results with large sample sizes.
In certain situations that shall be undoubtedly more and more common in the Big Data era, the datasets available are so massive that computing statistics over the full sample is hardly feasible, if not unfeasible. A natural approach in this context consists in using survey schemes and substituting the "full data" stati…
Hexagonal tilings minimize perimeter with unequal volumes.
A new method approximates deep neural networks using Kalman Filters.
New algorithm clusters Gaussian mixtures with unknown covariance efficiently.
The covering spectrum is a geometric invariant of a Riemannian manifold, more generally of a metric space, that measures the size of its one-dimensional holes by isolating a portion of the length spectrum. In a previous paper we demonstrated that the covering spectrum is not a spectral invariant of a manifold in dimens…
A new classification rule for FDA improves classification performance by accounting for unequal covariance matrices.
Combines cost-sensitive and Neyman-Pearson paradigms for better binary classification.
We study the misclassification error for community detection in general heterogeneous stochastic block models (SBM) with noisy or partial label information. We establish a connection between the misclassification rate and the notion of minimum energy on the local neighborhood of the SBM. We develop an optimally weighte…
Optimal testing of discrete distributions with high probability, achieving sample complexity bounds.
Importance sampling is often used in machine learning when training and testing data come from different distributions. In this paper we propose a new variant of importance sampling that can reduce the variance of importance sampling-based estimates by orders of magnitude when the supports of the training and testing d…
IMPACT optimizes LLM compression by focusing on activation importance, reducing model size up to 55.4%.
This study examines the topology of singularities in optimal semicouplings between unequal spaces.
Graph clustering involves the task of dividing nodes into clusters, so that the edge density is higher within clusters as opposed to across clusters. A natural, classic and popular statistical setting for evaluating solutions to this problem is the stochastic block model, also referred to as the planted partition model…
Recently, several studies have explored methods for using KG embedding to answer logical queries. These approaches either treat embedding learning and query answering as two separated learning tasks, or fail to deal with the variability of contributions from different query paths. We proposed to leverage a graph attent…
The MobileNet model was used by applying transfer learning on the 7 skin diseases to create a skin disease classification system on Android application. The proponents gathered a total of 3,406 images and it is considered as imbalanced dataset because of the unequal number of images on its classes. Using different samp…
Research into time series classification has tended to focus on the case of series of uniform length. However, it is common for real-world time series data to have unequal lengths. Differing time series lengths may arise from a number of fundamentally different mechanisms. In this work, we identify and evaluate two cla…
Paper tackles robust estimation of tree-structured Ising models without side information.
A statistical generalization is made of microeconomics in the spirit of going from classical to statistical mechanics. The price and quantity of every commodity1 traded in the market, at each instant of time, is considered to be an independent random variable: all prices and quantities are considered to be stochastic p…
Developing tools for computing string amplitudes with hyperbolic vertices.
Recent work shows unequal performance of commercial face classification services in the gender classification task across intersectional groups defined by skin type and gender. Accuracy on dark-skinned females is significantly worse than on any other group. In this paper, we conduct several analyses to try to uncover t…
The twentieth century was a period of outstanding economic growth together with an unequal income distribution. This paper analyses the international distribution of growth rates and its dynamics during the twentieth century. We show that the whole century is characterized by a high heterogeneity in the distribution of…
Cyclical MCMC tackles high-dimensional multimodal distributions, showing convergence under certain conditions.
New method uses causal thinking to make AI fairer decisions.
Study shows how learning and analytical models affect reneging and jockeying in a dual M/M/1 system.
The paper analyzes trade dynamics among G7 countries, revealing unequal exchange and degenerate equilibrium states.
Combines BART and Gaussian process for spatial covariate prediction with uncertainty.
Sharp spectral theorems and isoperimetric inequalities for manifolds with nonnegative Ricci curvature.
New framework controls statistical dispersion for high-stakes applications.
This paper introduces forward-looking measures of the network connectedness of fears in the financial system, arising due to the good and bad beliefs of market participants about uncertainty that spreads unequally across a network of banks. We argue that this asymmetric network structure extracted from call and put tra…
A new portfolio method using quantum mechanics improves risk diversification.
Robust fuzzy clustering for EEG driver alertness with outlier detection.
Learnable multiclass hypothesis classes don't always have a sample compression scheme of fixed size.
In biospectroscopy, suitably annotated and statistically independent samples (e. g. patients, batches, etc.) for classifier training and testing are scarce and costly. Learning curves show the model performance as function of the training sample size and can help to determine the sample size needed to train good classi…
Median-of-means sampling outperforms mean-of-means for large sample sizes in numerical integration.
Study examines how taxes affect wealth inequality in economic models.
This paper studies the problem of finding the exact ranking from noisy comparisons. A comparison over a set of items produces a noisy outcome about the most preferred item, and reveals some information about the ranking. By repeatedly and adaptively choosing items to compare, we want to fully rank the items with a …
Controversies around race and machine learning have sparked debate among computer scientists over how to design machine learning systems that guarantee fairness. These debates rarely engage with how racial identity is embedded in our social experience, making for sociological and psychological complexity. This complexi…
New convergence results for NGVI with various step sizes and sample sizes.
Adapts DR objectives for both sample and feature size reduction.
We obtain the first positive results for bounded sample compression in the agnostic regression setting with the loss, where . We construct a generic approximate sample compression scheme for real-valued function classes exhibiting exponential size in the fat-shattering dimension but independen…
pmsims R package uses Gaussian process for flexible sample size estimation in clinical models.
Leveraging reference-only samples for two-sample testing under size asymmetry
Bayesian approach learns linear networks from high-dimensional data.