Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

2875738601,146 · Jun 202019922001200920172026
48 results for data uniformity

New approach to adversarial robustness with non-uniform perturbations.

problem Real-world adversaries craft adversarial examples with non-uniform perturbations.
method Proposes non-uniform perturbations based on feature dependencies and data distribution.
result Shows improved robustness to real-world attacks compared to uniform perturbations.

PAC-Bayesian theory applied to data-dependent hypothesis sets yields uniform generalization bounds.

problem Proving uniform generalization bounds for data-dependent hypothesis sets.
method Applying PAC-Bayesian framework on 'random sets' and considering data-dependent hypothesis sets.
result Data-dependent uniform generalization bounds are proven, providing tighter and unified results.

This work studies the robustness certification problem of neural network models, which aims to find certified adversary-free regions as large as possible around data points. In contrast to the existing approaches that seek regions bounded uniformly along all input features, we consider non-uniform bounds and use it to …

2019-03-15abs ↗pdf ↗

The paper explores uniform perfectness and centers in Morse boundaries.

problem Detecting κκ-center exhaustivity in uniformly perfect Morse boundaries.
method Analyzes CAT(0) and geodesic spaces, using visual boundary data and metric transforms.
result Fixed-basepoint uniform perfectness is insufficient for κκ-center exhaustivity.

Deep learning generalizes well despite being overparameterized.

problem Why deep networks generalize well despite fitting training data perfectly.
method Empirical study of training methods and derivation of data-dependent generalization bounds.
result Uniform convergence alone is insufficient for explaining generalization in overparameterized settings.

The paper proposes a uniformity regularization scheme to improve deep neural network transferability.

problem Improving deep neural network transferability and adaptation to new tasks.
method Introduces a uniformity regularization scheme to encourage high uniformity in embedding space.
result Uniformity regularization consistently offers benefits over baseline methods and achieves state-of-the-art performance in Deep Metric Learning and Meta-Learning.

The paper classifies test configurations and derives a criterion for uniform K-stability of certain algebraic varieties.

problem Uniform K-stability of GG-varieties of complexity 1.
method Classification of GG-equivariant normal test configurations via combinatorial data and derivation of a criterion for uniform K-stability.
result Derivation of a criterion for uniform K-stability in terms of combinatorial data.

Frequency bias affects neural network training on non-uniform data.

problem Understanding how frequency bias impacts neural networks trained on non-uniformly distributed data.
method Used the Neural Tangent Kernel (NTK) model to explore the effect of variable density on training dynamics.
result Convergence time for learning a pure harmonic function depends on the local density at a point.

We design and mathematically analyze sampling-based algorithms for regularized loss minimization problems that are implementable in popular computational models for large data, in which the access to the data is restricted in some way. Our main result is that if the regularizer's effect does not become negligible as th…

2019-05-26abs ↗pdf ↗

The paper proposes a method for generating uniform interpolations on data manifolds.

problem Generating high-quality interpolations between data samples on complex manifolds.
method Autoencoder network with interpolation network, regularized by a Riemannian metric.
result The method generates interpolations that remain within the manifold's distribution.

A new UU-test decides unimodality of datasets.

problem Deciding on the unimodality of a dataset for better data analysis.
method UU-test operates on the empirical cumulative density function (ecdf) to build a piecewise linear approximation that models the data as a Uniform Mixture Model.
result The UU-test provides a statistical model of the data in the form of a Uniform Mixture Model.

A new model-free subsampling method using uniform designs is proposed.

problem Model-based subsampling methods are often dependent on model assumptions.
method Developed a criterion (GEFD) and a model-free subsampling method based on uniform designs.
result The proposed method outperforms random sampling and is robust under diverse model specifications.

Bayesian network structure learning is often performed in a Bayesian setting, evaluating candidate structures using their posterior probabilities for a given data set. Score-based algorithms then use those posterior probabilities as an objective function and return the maximum a posteriori network as the learned model.…

2017-04-12abs ↗pdf ↗

Paper tackles estimating initial conditions of spatio-temporal processes from sparse data.

problem Estimating initial conditions of spatio-temporal advection-diffusion processes from sparse data.
method Regularized convex optimization problem with Alternating Direction Method of Multipliers.
result Efficient solutions for non-uniform and shifted uniform sampling schemes.

Study tests uniformity of categorical data against missing-ball alternatives, finding chi-squared test outperforms.

problem Testing uniformity of categorical data against missing-ball alternatives.
method Characterizes minimax risk, uses collisions and chi-squared test, reduces to structured subset of alternatives.
result Minimax test outperforms chi-squared test under least favorable alternative.

The paper explores why a specific type of predictor works well in noisy data.

problem Understanding why a specific type of predictor (minimum-norm interpolator) works well in noisy data.
method The paper uses uniform convergence and zero-error predictors in a norm ball to explain the success of the minimum-norm interpolator.
result The minimum-norm interpolator is consistent, and this can be explained by uniform convergence of zero-error predictors in a norm ball.

We revisit Spakula's uniform K-homology, construct the external product for it and use this to deduce homotopy invariance of uniform K-homology. We define uniform K-theory and on manifolds of bounded geometry we give an interpretation of it via vector bundles of bounded geometry. We further construct a cap product with…

2018-08-23abs ↗pdf ↗

Random sampling has become a critical tool in solving massive matrix problems. For linear regression, a small, manageable set of data rows can be randomly selected to approximate a tall, skinny data matrix, improving processing time significantly. For theoretical performance guarantees, each row must be sampled with pr…

2014-08-21abs ↗pdf ↗

Improved sampling accuracy in SG-MCMC methods via non-uniform gradient subsampling.

problem Computational inefficiency and sampling error in stochastic gradient MCMC methods.
method Proposes a non-uniform subsampling scheme to reduce sampling error in EWSG, a variant of SG-MCMC.
result EWSG reduces sampling error compared to uniform subsampling, improving accuracy without sacrificing convergence speed.

In "Rips complexes and covers in the uniform category" \cite{Rips} the authors define, following James \cite{J}, covering maps of uniform spaces and introduce the concept of generalized uniform covering maps. Conditions for the existence of universal uniform covering maps and generalized uniform covering maps are given…

2010-08-02abs ↗pdf ↗

Batch normalization biases linear models towards uniform margins, improving performance in binary classification.

problem Understanding the implicit bias of batch normalization in linear models and neural networks.
method Analyzing gradient descent convergence on linear models and two-layer CNNs with batch normalization.
result Gradient descent with batch normalization in linear models converges to a uniform margin classifier with an exponential convergence rate.

This paper refines homotopy theory for cubical sets and uniform spaces.

problem Classical homotopy theory limitations in cubical sets and uniform spaces.
method Develops a uniform-theoretic refinement for cubical sets and uniform spaces, lifting to a full and faithful embedding.
result Lifts classical homotopy categories to new uniform homotopy categories, generalizing cohomology theories.

Validates conformal prediction for network data under non-uniform sampling.

problem Validity of conformal prediction for network data under non-representative sampling.
method Interprets sampling mechanisms as selection rules, studies validity conditional on selection events, uses permutation invariance and joint exchangeability.
result Finite-sample validity of conformal prediction for certain selection events and asymptotic validity for random walk sampling.

Study shows gap between uniform convergence and test error in random feature models.

problem Understanding the gap between uniform convergence and test error in random feature models.
method Analytical expressions for uniform convergence over norm balls, interpolators, and minimum norm interpolator risk derived and proved.
result Uniform convergence over interpolators still gives a non-trivial bound of test error even when classical uniform convergence is vacuous.

New loss function equivalence reveals PER's uniform sampling can be improved.

problem Improving Prioritized Experience Replay (PER) for better learning efficiency.
method Transforming non-uniformly sampled data loss functions into uniformly sampled ones.
result Some environments can replace PER with a new loss function without performance loss.

Selecting more uniformly distributed data improves training efficiency and performance.

problem Improving data selection for training large language models (LLMs).
method Established a convergence framework for gradient descent beyond the NTK regime, proving that more uniform data leads to larger minimum pairwise distances and faster training.
result Selecting more uniformly distributed data accelerates training and achieves comparable or better performance in LLMs.

The paper provides bounds on the CDF of a variable under nonstationary conditions.

problem Estimating the complete distribution of a random variable under nonstationary conditions.
method Time-uniform and value-uniform bounds on the CDF of the running averaged conditional distribution.
result Presented computationally efficient bounds that are always valid and sometimes trivial.

The theory of learning under the uniform distribution is rich and deep, with connections to cryptography, computational complexity, and the analysis of boolean functions to name a few areas. This theory however is very limited due to the fact that the uniform distribution and the corresponding Fourier basis are rarely …

2013-07-13abs ↗pdf ↗