Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

166333499665 · Jun 202019922001200920172026
48 results for Unequal Sample Sizes

New method for MMD with unequal sample sizes improves test power.

problem Existing MMD methods assume equal sample sizes, discarding valuable data.
method Extended generalized U-statistics to handle unequal sample sizes.
result New asymptotic distributions and power optimization for MMD with unequal sample sizes.

We consider the problem of closeness testing for two discrete distributions in the practically relevant setting of \emph{unequal} sized samples drawn from each of them. Specifically, given a target error parameter ε>0\varepsilon > 0, m1m_1 independent draws from an unknown distribution p,p, and m2m_2 draws from an unkno…

2015-04-17abs ↗pdf ↗

The paper improves A/B testing for non-Gaussian data, ensuring reliable results with large sample sizes.

problem Inaccurate A/B testing results due to non-normal data and unequal sample sizes.
method Derives explicit formulas for minimum sample size and introduces an Edgeworth-based correction.
result Corrected method improves reliability of A/B testing in real-world conditions.

The covering spectrum is a geometric invariant of a Riemannian manifold, more generally of a metric space, that measures the size of its one-dimensional holes by isolating a portion of the length spectrum. In a previous paper we demonstrated that the covering spectrum is not a spectral invariant of a manifold in dimens…

2010-06-28abs ↗pdf ↗

A new classification rule for FDA improves classification performance by accounting for unequal covariance matrices.

problem Unequal covariance matrices in practical situations affect the performance of FDA and its variants.
method Proposes a novel classification rule for FDA that accounts for unequal covariance matrices, applicable to many FDA variants.
result The new classification rule improves classification performance compared to original FDA and variants.

Combines cost-sensitive and Neyman-Pearson paradigms for better binary classification.

problem Asymmetric binary classification problems with unequal error severities.
method Develops TUBE-CS algorithm to bridge cost-sensitive and Neyman-Pearson paradigms.
result High-probability control of population type I error.

Optimal testing of discrete distributions with high probability, achieving sample complexity bounds.

problem Testing discrete distributions with high probability accuracy.
method Characterizing sample complexity as a function of parameters like δ, providing sample-optimal testers.
result Optimal algorithms for closeness and independence testing, achieving within constant factors of information-theoretic lower bounds.

Importance sampling is often used in machine learning when training and testing data come from different distributions. In this paper we propose a new variant of importance sampling that can reduce the variance of importance sampling-based estimates by orders of magnitude when the supports of the training and testing d…

2016-11-10abs ↗pdf ↗

IMPACT optimizes LLM compression by focusing on activation importance, reducing model size up to 55.4%.

problem Resource constraints in deploying large language models (LLMs).
method IMPACT integrates activation importance into low-rank compression, optimizing for both size and accuracy.
result IMPACT achieves up to 55.4% greater model size reduction while maintaining comparable or better accuracy.

This study examines the topology of singularities in optimal semicouplings between unequal spaces.

problem Topology of singularities in optimal semicouplings between unequal spaces.
method Continuous strong deformation retracts and Uniform Halfspace condition.
result Homotopy-reductions from a source space onto singularities of cc-optimal semicouplings.

Graph clustering involves the task of dividing nodes into clusters, so that the edge density is higher within clusters as opposed to across clusters. A natural, classic and popular statistical setting for evaluating solutions to this problem is the stochastic block model, also referred to as the planted partition model…

2012-10-11abs ↗pdf ↗

The MobileNet model was used by applying transfer learning on the 7 skin diseases to create a skin disease classification system on Android application. The proponents gathered a total of 3,406 images and it is considered as imbalanced dataset because of the unequal number of images on its classes. Using different samp…

2019-11-13abs ↗pdf ↗

Research into time series classification has tended to focus on the case of series of uniform length. However, it is common for real-world time series data to have unequal lengths. Differing time series lengths may arise from a number of fundamentally different mechanisms. In this work, we identify and evaluate two cla…

2019-10-10abs ↗pdf ↗

Paper tackles robust estimation of tree-structured Ising models without side information.

problem Learning tree-structured Ising models with flipped signs of variables.
method Proves unidentifiability, proposes an algorithm with logarithmic sample complexity and polynomial run-time complexity.
result Empirically demonstrates robustness of proposed algorithm in the flipped signs setting.

A statistical generalization is made of microeconomics in the spirit of going from classical to statistical mechanics. The price and quantity of every commodity1 traded in the market, at each instant of time, is considered to be an independent random variable: all prices and quantities are considered to be stochastic p…

2012-11-30abs ↗pdf ↗

Recent work shows unequal performance of commercial face classification services in the gender classification task across intersectional groups defined by skin type and gender. Accuracy on dark-skinned females is significantly worse than on any other group. In this paper, we conduct several analyses to try to uncover t…

2018-11-30abs ↗pdf ↗

The twentieth century was a period of outstanding economic growth together with an unequal income distribution. This paper analyses the international distribution of growth rates and its dynamics during the twentieth century. We show that the whole century is characterized by a high heterogeneity in the distribution of…

2017-08-22abs ↗pdf ↗

Cyclical MCMC tackles high-dimensional multimodal distributions, showing convergence under certain conditions.

problem High-dimensional multimodal posterior distributions in deep learning.
method Cyclical MCMC framework that tracks tempered versions of the target distribution over time.
result Cyclical MCMC converges to the target distribution under fast mixing kernels but fails in slow mixing cases.

New method uses causal thinking to make AI fairer decisions.

problem Designing fair machine learning models that treat equal individuals equally and unequals unequally.
method Rank-preserving interventional distributions and warping method.
result Warping method effectively identifies discriminated individuals and mitigates unfairness.

Study shows how learning and analytical models affect reneging and jockeying in a dual M/M/1 system.

problem How do learning and analytical models affect reneging and jockeying in a dual M/M/1 system?
method Analytical and online trained actor-critic models were used to study reneging and jockeying in a dual M/M/1 system.
result Both analytical and online trained actor-critic models yield the same asymptotic limits for reneging and jockeying, but differ in practical sizes.

The paper analyzes trade dynamics among G7 countries, revealing unequal exchange and degenerate equilibrium states.

problem Unequal exchange and degenerate equilibrium states in international trade among G7 countries.
method Analysis based on a model of international trade with supply and demand structures.
result Found relative equilibrium price vector is very degenerate, indicating unequal exchange.

Combines BART and Gaussian process for spatial covariate prediction with uncertainty.

problem Improving spatial prediction models with nonlinear and interaction covariates.
method Bayesian Additive Regression Trees (BART) combined with Gaussian process for spatial dependence.
result Effective in reducing computational burden through INLA and MCMC.

Sharp spectral theorems and isoperimetric inequalities for manifolds with nonnegative Ricci curvature.

problem Understanding the geometry and topology of manifolds with nonnegative Ricci curvature.
method New spectral inequalities and isoperimetric problems involving unequal weights and warped bubbles.
result Sharp spectral and isoperimetric bounds for manifolds with nonnegative Ricci curvature.

New framework controls statistical dispersion for high-stakes applications.

problem Understanding and controlling the dispersion of loss distributions in high-stakes applications.
method Simple yet flexible framework for distribution-free control of statistical dispersion measures.
result Proposed methods control statistical dispersion measures with societal implications.

This paper introduces forward-looking measures of the network connectedness of fears in the financial system, arising due to the good and bad beliefs of market participants about uncertainty that spreads unequally across a network of banks. We argue that this asymmetric network structure extracted from call and put tra…

2018-10-29abs ↗pdf ↗

A new portfolio method using quantum mechanics improves risk diversification.

problem Improving risk-based portfolio construction methods for multi-asset portfolios.
method Schrödinger principal component analysis applied to extract common factors from asset fluctuations.
result The proposed method outperforms conventional risk parity and other risk diversification methods.

Robust fuzzy clustering for EEG driver alertness with outlier detection.

problem Ambiguous state boundaries in multivariate time series data.
method RFCPCA, a robust fuzzy subspace-clustering method for MTS.
result RFCPCA improves clustering accuracy and characterizes uncertainty and outliers in MTS.

Learnable multiclass hypothesis classes don't always have a sample compression scheme of fixed size.

problem The limitation of sample compression schemes for multiclass hypothesis classes.
method Analysis of DS dimension and sample compression schemes.
result Learnable multiclass hypothesis classes do not always have a sample compression scheme of fixed size.

In biospectroscopy, suitably annotated and statistically independent samples (e. g. patients, batches, etc.) for classifier training and testing are scarce and costly. Learning curves show the model performance as function of the training sample size and can help to determine the sample size needed to train good classi…

2012-11-06abs ↗pdf ↗

Median-of-means sampling outperforms mean-of-means for large sample sizes in numerical integration.

problem Improving numerical integration accuracy in high dimensions.
method Median-of-means sampling compared to mean-of-means using RQMC methods.
result Median-of-means sampling is superior for large sample sizes, while mean-of-means is better for smaller sample sizes.

Controversies around race and machine learning have sparked debate among computer scientists over how to design machine learning systems that guarantee fairness. These debates rarely engage with how racial identity is embedded in our social experience, making for sociological and psychological complexity. This complexi…

2018-11-28abs ↗pdf ↗

New convergence results for NGVI with various step sizes and sample sizes.

problem Understanding convergence of stochastic NGVI for various schedules.
method Projected stochastic NGVI for exponential family variational distributions.
result Geometric convergence and $\mathcal{O}\left(\frac{1}{T^ρ} ight)$ rates for different schedules.

We obtain the first positive results for bounded sample compression in the agnostic regression setting with the p\ell_p loss, where p[1,]p\in [1,\infty]. We construct a generic approximate sample compression scheme for real-valued function classes exhibiting exponential size in the fat-shattering dimension but independen…

2018-10-03abs ↗pdf ↗

pmsims R package uses Gaussian process for flexible sample size estimation in clinical models.

problem Determining adequate sample size for clinical prediction models.
method Simulation-based Gaussian process search for flexible sample size estimation.
result Gaussian process-based method produces more stable sample size estimates, especially in challenging settings.

Bayesian approach learns linear networks from high-dimensional data.

problem Learning high-dimensional linear Bayesian networks.
method Iterative estimation of topological ordering and parents using inverse partial covariance matrix with Bayesian regularization.
result The method successfully recovers network structure under certain conditions.