Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

72144215287 · Jun 202019922001200920172026
48 results for super-polynomial convergence

The study reveals the efficiency of sampling from tilted distributions.

problem Sampling from a tilted distribution of an unknown underlying distribution.
method Self-normalized importance sampling to characterize accuracy.
result Polynomial vs super-polynomial sample complexity for bounded vs unbounded distributions.

We provide a detailed study on the implicit bias of gradient descent when optimizing loss functions with strictly monotone tails, such as the logistic loss, over separable datasets. We look at two basic questions: (a) what are the conditions on the tail of the loss function under which gradient descent converges in the…

2018-03-05abs ↗pdf ↗

In this article we construct closed, isospectral, non-isometric locally symmetric manifolds. We have three main results. First, we construct arbitrarily large sets of closed, isospectral, non-isometric manifolds. Second, we show the growth of size these sets of isospectral manifolds as a function of volume is super-pol…

2006-06-21abs ↗pdf ↗

We explain how to adapt a construction of M. Sageev's to construct a proper action on a CAT(0) cube complex starting from a proper action on a wall space, and use this to deduce that if G is a group containing an amenable subgroup H of super-polynomial growth and G acts properly on a space with walls then there are arb…

2003-09-02abs ↗pdf ↗

We prove that the colored HOMFLY polynomial of a link, colored by symmetric or exterior powers of the fundamental representation, is q-holonomic with respect to the color parameters. As a result, we obtain the existence of an (a,q) super-polynomial of all knots in 3-space. Our result has implications on the quantizatio…

2012-11-27abs ↗pdf ↗

In many estimation problems, e.g. linear and logistic regression, we wish to minimize an unknown objective given only unbiased samples of the objective function. Furthermore, we aim to achieve this using as few samples as possible. In the absence of computational constraints, the minimizer of a sample average of observ…

2014-12-20abs ↗pdf ↗

We prove that the HOMFLYPT polynomial of a link, colored by partitions with a fixed number of rows is a qq-holonomic function. Specializing to the case of knots colored by a partition with a single row, it proves the existence of an (a,q)(a,q) super-polynomial of knots in 3-space, as was conjectured by string theorists. …

2016-04-28abs ↗pdf ↗

Enhances quantum computing for symmetrical systems, proving a new class of problems.

problem Proving the efficiency of a new quantum computing model for symmetrical systems.
method Introducing equivariant convolutional quantum algorithms tailored for SU(d) symmetries.
result Demonstrates a problem that can be solved efficiently on a new quantum model, suggesting it's not classically simulatable.

We modify our previous construction of link homology in order to include a natural duality functor F\mathfrak{F}. To a link LL we associate a triply-graded module HXY(L)HXY(L) over the graded polynomial ring R(L)=C[x1,y1,,x,y]R(L)=\mathbb{C}[x_1,y_1,\dots,x_\ell,y_\ell]. The module has an involution F\mathfrak{F} that intertwines the F…

2019-05-16abs ↗pdf ↗

New algorithm learns ReLU networks efficiently using Schur polynomials.

problem PAC learning a linear combination of ReLU activations under Gaussian distribution.
method Uses tensor decomposition and Schur polynomials to identify and analyze higher-order moments.
result Near-optimal sample and computational complexity for learning ReLU networks.

WildCat efficiently compresses neural network attention mechanisms.

problem Expensive quadratic runtime of attention mechanisms in neural networks.
method Uses a weighted coreset with a fast subsampling algorithm to approximate attention with near-linear time complexity.
result Approximates exact attention with super-polynomial error decay and near-linear runtime.

The paper proves barriers to approximating functions with small weights and depth in neural networks.

problem Proving barriers to approximating functions with constant depth neural networks.
method Reduction to open problems and natural-proof barriers in circuit complexity, and a new approach to polynomially-bounded functions.
result There are fundamental barriers to proving results beyond depth 4 for constant-depth neural networks.

Modern inference and learning often hinge on identifying low-dimensional structures that approximate large scale data. Subspace clustering achieves this through a union of linear subspaces. However, in contemporary applications data is increasingly often incomplete, rendering standard (full-data) methods inapplicable. …

2018-08-02abs ↗pdf ↗

New framework formalizes RLHF trilemma: improving safety, fairness, and robustness is computationally infeasible.

problem Aligning large language models with diverse human values while maintaining computational feasibility and robustness.
method Complexity-theoretic analysis integrating statistical learning theory and robust optimization.
result Achieving both representativeness (epsilon <= 0.01) and robustness (delta <= 0.001) for global-scale populations requires super-polynomial operations.

This paper strengthens the computational separation between multimodal and unimodal learning, showing unimodal learning is hard on typical instances.

problem Theoretical justification for empirical success of multimodal machine learning.
method Introduced a stronger average-case computational separation between unimodal and multimodal learning.
result For typical instances, unimodal learning is computationally hard, while multimodal learning is easy.

The paper explores how LLMs with CoT improve performance on complex tasks.

problem Understanding the mechanisms behind LLMs' improved performance with CoT.
method Using circuit complexity theory, the paper examines LLMs' expressivity in solving mathematical and decision-making problems.
result LLMs with CoT can generate correct solutions step-by-step, even for complex tasks.

The paper tackles robust policy learning in multitask contextual bandits with adversarial users.

problem Learning optimal policies in multitask contextual bandits with a small fraction of adversarial users.
method Developed efficient robust mean estimators for both uni-variate and high-dimensional random variables.
result Lower bound of ildeΩ(min(S,A)α2/ε2) ildeΩ(\min(S,A) \cdot α^2 / ε^2) per-user interactions to learn an εε-optimal policy for good users.

Algorithm learns near-optimal policies for reward-mixing MDPs with few latent contexts.

problem Episodic reinforcement learning in reward-mixing Markov decision processes with a few latent contexts.
method Sample-efficient algorithm EM^2 using higher-order method-of-moments approach.
result Provides an ε-optimal policy using O(ε^(-2) * S^d A^d * poly(H, Z)^d) episodes for arbitrary M ≥ 2.

Efficient algorithm for online learning with Massart noise achieves near-optimal mistake bound.

problem Online learning with adversarial context and Massart noise.
method Developed an efficient algorithm for γγ-margin linear classifiers in the presence of Massart noise.
result Achieved a mistake bound of ηT+o(T)ηT + o(T) for the online learning model.

Polynomial-time algorithm learns high-dimensional halfspaces without labels.

problem Learning high-dimensional halfspaces with margins in polynomial time.
method Contrastive moments and polynomial-time algorithm.
result Establishes the unique and efficient identifiability of the hidden halfspace.

The goal of this paper is to characterize function distributions that deep learning can or cannot learn in poly-time. A universality result is proved for SGD-based deep learning and a non-universality result is proved for GD-based deep learning; this also gives a separation between SGD-based deep learning and statistic…

2020-01-07abs ↗pdf ↗

This paper is concerned with jointly recovering nn node-variables {xi}1in\left\{ x_{i}\right\}_{1\leq i\leq n} from a collection of pairwise difference measurements. Imagine we acquire a few observations taking the form of xixjx_{i}-x_{j}; the observation pattern is represented by a measurement graph G\mathcal{G} with an ed…

2015-04-06abs ↗pdf ↗

As the success of deep learning reaches more grounds, one would like to also envision the potential limits of deep learning. This paper gives a first set of results proving that certain deep learning algorithms fail at learning certain efficiently learnable functions. The results put forward a notion of cross-predictab…

2018-12-16abs ↗pdf ↗

Two algorithms learn Gaussian graphical models from Glauber dynamics trajectories, achieving optimal performance.

problem Learning Gaussian graphical models from a single trajectory of a dependent stochastic process.
method Two algorithms based on dueling-neighborhood search and local statistics built from the update sequence of Glauber dynamics.
result Achieve κ2κ^{-2} dependence of the information-theoretic lower bounds, mixing-free and signal-optimal.

Paper addresses classification under misspecification, providing simpler algorithms and resolving open questions.

problem Learning halfspaces and generalized linear models under Massart noise and other corruption models.
method Developed simpler algorithms and used blackbox knowledge distillation to convert complex classifiers to proper ones. Leveraged evolvability for theoretical insights.
result First efficient algorithm for learning Massart halfspaces with η+εη+ ε accuracy, and general algorithm for generalized linear models.

New lower bounds show learning intersections of halfspaces is hard even for a few halfspaces.

problem Learning intersections of halfspaces in polynomial time under standard assumptions.
method Unified connection to parallel pancakes distribution for proving hardness.
result Learning ω(loglogN)ω(\log \log N) halfspaces in dimension NN requires super-polynomial time under standard assumptions.

Efficient algorithm learns mixture models of heavy-tailed distributions.

problem Learning mixture models of heavy-tailed distributions.
method Efficient high-dimensional sparse Fourier transforms.
result Algorithm succeeds for heavy-tailed distributions, including Laplace but excluding Gaussians.

This is an intuitive survey of extrinsic and intrinsic notions of convergence of manifolds complete with pictures of key examples and a discussion of the properties associated with each notion. We begin with a description of three extrinsic notions which have been applied to study sequences of submanifolds in Euclidean…

2010-06-02abs ↗pdf ↗

The abstract discusses convergence properties of Lipschitz functions and sets defined by equations.

problem Convergence of Lipschitz functions and sets defined by equations.
method Painlevé-Kuratowski convergence applied to Lipschitz functions and sets defined by equations.
result Generalizations and reverses of classical theorems on convergence of functions and sets.

The objective of this paper is to introduce the notion of generalized almost statistical (briefly, GAS) convergence of bounded real sequences, which generalizes the notion of almost convergence as well as statistical convergence of bounded real sequences. As a special kind of Banach limit functional, we also introduce …

2019-11-15abs ↗pdf ↗