Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Dec 199219922001200920172026
48 results for minimal size

SGD's performance improves with critical batch size, minimizing SFO complexity.

problem Optimizing SGD's performance with batch size and learning rate.
method Analysis of SGD using constant and decaying learning rates, focusing on batch size effects.
result SGD with critical batch size minimizes SFO complexity.

DPSM minimizes prediction set size by integrating conformal principles into deep classifier training.

problem Large prediction sets from standard conformal methods are impractical.
method Formulates conformal training as bilevel optimization, proposing DPSM algorithm.
result Significantly reduces prediction set size compared to prior methods.

We speed up marginal inference by ignoring factors that do not significantly contribute to overall accuracy. In order to pick a suitable subset of factors to ignore, we propose three schemes: minimizing the number of model factors under a bound on the KL divergence between pruned and full models; minimizing the KL dive…

2012-03-15abs ↗pdf ↗

We study the Stochastic Gradient Descent (SGD) method in nonconvex optimization problems from the point of view of approximating diffusion processes. We prove rigorously that the diffusion process can approximate the SGD algorithm weakly using the weak form of master equation for probability evolution. In the small ste…

2017-05-22abs ↗pdf ↗

Optimal batch size minimizes training time for neural networks.

problem Minimizing training time for two-layer neural networks with SGD.
method Characterized optimal batch size as a function of target hardness (information exponents). Used Correlation loss SGD to overcome limitations.
result Optimal batch size minimizes training time without changing total sample complexity.

Permutation of any two hidden units yields invariant properties in typical deep generative neural networks. This permutation symmetry plays an important role in understanding the computation performance of a broad class of neural networks with two or more hidden units. However, a theoretical study of the permutation sy…

2019-04-30abs ↗pdf ↗

This is the third paper in a series devoted to enumerating the prime alternating knots and links. This paper establishes a method for enumerating the prime alternating links. It is shown that one may choose any prime alternating link diagram of a given minimal crossing size and by applications of just two operators (T …

2002-11-28abs ↗pdf ↗

We prove that the rank (that is, the minimal size of a generating set) of lattices in a general connected Lie group is bounded by the co-volume of the projection of the lattice to the semi-simple part of the group. This was proved by Gelander for semi-simple Lie groups and by Mostow for solvable Lie groups. Here we con…

2019-03-12abs ↗pdf ↗

Let FgF_g be a closed orientable surface of genus gg. A set Ω={γ1,,γs}Ω= \{ γ_1, \dots, γ_s\} of pairwise non-homotopic simple closed curves on FgF_g is called a \emph{filling system} or simply a \emph{filling} of FgF_g, if FgΩF_g\setminus Ω is a union of bb topological discs for some b1b\geq 1. A filling system is called \em…

2017-08-23abs ↗pdf ↗

New adaptive scheduler improves SAM for better model training.

problem Training machine learning models requires selecting a learning rate, which is often difficult and time-consuming.
method Derive Polyak schedulers tailored to SAM-style updates, proving linear convergence for strongly convex objectives and an O(1/T) rate for convex objectives.
result Polyak schedulers achieve comparable or better performance than tuned SAM baselines, reducing the need for learning-rate tuning.

New study shows ERMs can fail in convex optimization with high dimensionality.

problem The limitations of Empirical Risk Minimizer in high-dimensional stochastic convex optimization.
method Constructed a specific instance showing ERMs can be unique and overfit.
result Gradient Descent can also overfit in certain conditions, resolving a gap in lower bounds.

Classifies knots by lattice size, finding unknot ratios and crossing numbers.

problem Understanding the distribution of knots within different lattice sizes.
method Introduced a new knot classification by lattice size, analyzed ratios of unknots and knots with more than 10 crossings, and compared with theoretical estimates.
result Ratio of unknots decreases exponentially with lattice size, and computational results match theoretical estimates.

We consider the problem of exact recovery of any m×nm\times n matrix of rank ϱ\varrho from a small number of observed entries via the standard nuclear norm minimization framework. Such low-rank matrices have degrees of freedom (m+n)ϱϱ2(m+n)\varrho - \varrho^2. We show that any arbitrary low-rank matrices can be recovered exa…

2015-03-22abs ↗pdf ↗

Improved Frank-Wolfe method reduces dependence on data size for empirical risk minimization.

problem Reducing dependence on number of data observations in Frank-Wolfe methods.
method Taylor-series approximated gradients applied to Frank-Wolfe method.
result Significant speed-ups over existing methods on real-world datasets.

Scaling laws for neural language models reveal optimal model size and compute allocation.

problem Understanding the optimal model size and compute allocation for neural language models.
method Empirical analysis of scaling laws for cross-entropy loss across model size, dataset size, and compute.
result Simple equations govern the dependence of overfitting and training speed on model/dataset size and model size, respectively.

Mini-batch stochastic gradient descent and variants thereof have become standard for large-scale empirical risk minimization like the training of neural networks. These methods are usually used with a constant batch size chosen by simple empirical inspection. The batch size significantly influences the behavior of the …

2016-12-15abs ↗pdf ↗

Kernel adaptive filters (KAF) are a class of powerful nonlinear filters developed in Reproducing Kernel Hilbert Space (RKHS). The Gaussian kernel is usually the default kernel in KAF algorithms, but selecting the proper kernel size (bandwidth) is still an open important issue especially for learning with small sample s…

2014-01-23abs ↗pdf ↗

The braid group's commutator subgroup is generated by two elements for n ≥ 7.

problem Generating the smallest possible generating sets for the commutator subgroup of braid groups.
method Analyzing specific cases of braid groups (n=4, 6, 5, 7+) to find generating sets of minimal size.
result For n ≥ 7, the commutator subgroup of the braid group is generated by two elements.

IMPACT optimizes LLM compression by focusing on activation importance, reducing model size up to 55.4%.

problem Resource constraints in deploying large language models (LLMs).
method IMPACT integrates activation importance into low-rank compression, optimizing for both size and accuracy.
result IMPACT achieves up to 55.4% greater model size reduction while maintaining comparable or better accuracy.

Research shows minimal communication limits adaptive function estimation rates.

problem Adaptive estimation of a smooth function under minimal communication constraints.
method Investigates the LL_\infty-risk and L2L_2-risk under different numbers of servers.
result For LL_\infty-risk, optimal rates cannot be achieved under minimal communication. For L2L_2-risk, adaptivity is possible but depends on server number and sample size.

Discrete approximation solves Björling's minimal surface problem.

problem Constructing minimal surfaces from real-analytic curves with specified normal fields.
method Approximate solution by discrete minimal surfaces and discrete isothermic surfaces.
result Approximation error is proportional to the square of the mesh size.

New method combines experimental and observational data for causal inference.

problem Combining internal validity of experiments and larger sample sizes of observations.
method Empirical risk minimization (ERM) framework with cross-validation.
result Efficacy and reliability demonstrated on real and synthetic data.

The variance reduction class of algorithms including the representative ones, SVRG and SARAH, have well documented merits for empirical risk minimization problems. However, they require grid search to tune parameters (step size and the number of iterations per inner loop) for optimal performance. This work introduces `…

2019-08-25abs ↗pdf ↗

Paper proposes a cost-sensitive conformal training method with provably controllable learning bounds.

problem Uncertainty quantification and learning bounds in conformal prediction.
method Cost-sensitive conformal training algorithm that minimizes the expected size of prediction sets using rank weighting.
result Theoretical analysis shows tightness between weighted objective and expected size of conformal prediction sets.

Gradient descent with large steps leads to chaotic parameter space and unpredictable outcomes.

problem Understanding the behavior of gradient descent with large step sizes in matrix factorization.
method Analyzing the fractal structure of the parameter space and deriving critical step sizes for convergence.
result Gradient descent with large steps exhibits chaotic behavior and sensitivity to initialization, creating a fractal boundary between converging and diverging minimizers.