Tiny benchmarks reduce LLM evaluation costs by using fewer examples.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Method trains deep models to explain predictions with fewer examples.
ActiveLab improves classifier accuracy with fewer annotations by re-labeling.
Machine learning models have been found to be susceptible to adversarial examples that are often indistinguishable from the original inputs. These adversarial examples are created by applying adversarial perturbations to input samples, which would cause them to be misclassified by the target models. Attacks that search…
One of the biggest bottlenecks in a machine learning workflow is waiting for models to train. Depending on the available computing resources, it can take days to weeks to train a neural network on a large dataset with many classes such as ImageNet. For researchers experimenting with new algorithmic approaches, this is …
Information planning enables faster learning with fewer training examples. It is particularly applicable when training examples are costly to obtain. This work examines the advantages of information planning for text data by focusing on three supervised models: Naive Bayes, supervised LDA and deep neural networks. We s…
Algorithm finds ribbon disks for alternating knots, resolving sliceness for most prime knots.
Several large classes of homogeneous spaces are known to be formal---in the sense of Rational Homotopy Theory. However, it seems that far fewer examples of non-formal homogeneous spaces are known. In this article we provide several construction principles and characterisations for non-formal homogeneous spaces, which w…
Regularization aims to improve prediction performance of a given statistical modeling approach by moving to a second approach which achieves worse training error but is expected to have fewer degrees of freedom, i.e., better agreement between training and prediction error. We show here, however, that this expected beha…
GGAN improves audio representation learning with fewer labels.
A new algorithm finds a separating hyperplane with fewer updates.
Error bounds based on worst likely assignments use permutation tests to validate classifiers. Worst likely assignments can produce effective bounds even for data sets with 100 or fewer training examples. This paper introduces a statistic for use in the permutation tests of worst likely assignments that improves error b…
Over the past year, the emergence of transfer learning with large-scale language models (LM) has led to dramatic performance improvements across a broad range of natural language understanding tasks. However, the size and memory footprint of these large LMs makes them difficult to deploy in many scenarios (e.g. on mobi…
The ability to learn from a small number of examples has been a difficult problem in machine learning since its inception. While methods have succeeded with large amounts of training data, research has been underway in how to accomplish similar performance with fewer examples, known as one-shot or more generally few-sh…
The purpose of this note is to define tri-moment maps for certain manifolds that carry closed non-degenerate 4-forms and an -action. Examples include quaternionic vector spaces and flag manifolds. We show how this map can be used ro reduce such manifolds to the ones with fewer symmetries. The images of such ma…
A core aspect of human intelligence is the ability to learn new tasks quickly and switch between them flexibly. Here, we describe a modular continual reinforcement learning paradigm inspired by these abilities. We first introduce a visual interaction environment that allows many types of tasks to be unified in a single…
We establish a correspondence between trisections of smooth, compact, oriented --manifolds with connected boundary and diagrams describing these trisected --manifolds. Such a diagram comes in the form of a compact, oriented surface with boundary together with three tuples of simple closed curves, with possibly fe…
We show how the success of deep learning could depend not only on mathematics but also on physics: although well-known mathematical theorems guarantee that neural networks can approximate arbitrary functions well, the class of functions of practical interest can frequently be approximated through "cheap learning" with …
We consider the "intrinsic" symmetry group of a two-component link , defined to be the image of the natural homomorphism from the standard symmetry group $\MCG(S^3,L)$ to the product $\MCG(S^3) \cross \MCG(L)$. This group, first defined by Whitten in 1969, records directly whether is isotopic to a link $L…
Computer program finds FAMED triangulations for thousands of knots.
Efficient kernel method learns differential equations with fewer data.
For a knot K, let b_n(K) be the minimum length of an n-stranded braid representative of K. Examples of knots exist for which b_n(K) is a non-increasing function. We investigate the behavior of b_n(K). We develop bounds on the function in terms of the genus of K, with stronger results for homogeneous knots and braid pos…
The Turaev genus and dealternating number of a link are two invariants that measure how far away a link is from alternating. We determine the Turaev genus of a torus knot with five or fewer strands either exactly or up to an error of at most one. We also determine the dealternating number of a torus knot with five or f…
Machine learning has made major advances in categorizing objects in images, yet the best algorithms miss important aspects of how people learn and think about categories. People can learn richer concepts from fewer examples, including causal models that explain how members of a category are formed. Here, we explore the…
In this paper, we define the primitive/Seifert-fibered property for a knot in S^3. If satisfied, the property ensures that the knot has a Dehn surgery that yields a small Seifert-fibered space (i.e. base S^2 and three or fewer critical fibers). Next we describe the twisted torus knots, which provide an abundance of exa…
Deep learning has achieved astonishing results on many tasks with large amounts of data and generalization within the proximity of training data. For many important real-world applications, these requirements are unfeasible and additional prior knowledge on the task domain is required to overcome the resulting problems…
We study the problem of generating adversarial examples in a black-box setting in which only loss-oracle access to a model is available. We introduce a framework that conceptually unifies much of the existing work on black-box attacks, and we demonstrate that the current state-of-the-art methods are optimal in a natura…
We show that there exist infinitely many examples of pairs of knots, K_1 and K_2, that have no epimorphism preserving peripheral structure although their A-polynomials have the factorization . Our construction accounts for most of the kno…
We show that the proportion of hyperbolic knots among all of the prime knots of or fewer crossings does not converge to as approaches infinity. Moreover, we show that if is a nontrivial knot then the proportion of satellites of among all of the prime knots of or fewer crossings does not converge…
Active feature selection uses mutual information to choose fewer labels for better feature selection.
The concordance orders of many algebraic order two knots of ten or fewer crossings have been heretofore unknown. We use Casson-Gordon invariants and twisted Alexander polynomials to find that, in all but one case, these knots do not have concordance order two. We also find that a certain family of algebraic order two t…
We investigate the bi-orderability of two-bridge knot groups and the groups of knots with 12 or fewer crossings by applying recent theorems of Chiswell, Glass and Wilson. Amongst all knots with 12 or fewer crossings (of which there are 2977), previous theorems were only able to determine bi-orderability of 599 of the c…
Note that this paper is superceded by "Black-Box Adversarial Attacks with Limited Queries and Information." Current neural network-based image classifiers are susceptible to adversarial examples, even in the black-box setting, where the attacker is limited to query access without access to gradients. Previous methods -…
We propose a hybrid approach aimed at improving the sample efficiency in goal-directed reinforcement learning. We do this via a two-step mechanism where firstly, we approximate a model from Model-Free reinforcement learning. Then, we leverage this approximate model along with a notion of reachability using Mean First P…
Most of Markov Chain Monte Carlo (MCMC) and sequential Monte Carlo (SMC) algorithms in existing probabilistic programming systems suboptimally use only model priors as proposal distributions. In this work, we describe an approach for training a discriminative model, namely a neural network, in order to approximate the …
S2cGAN uses fewer labels to train cGANs effectively.
We show that if is a nontrivial knot then the proportion of satellites of among all of the prime non-split links of or fewer crossings does not converge to as approaches infinity. This implies in particular that the proportion of hyperbolic links among all of the prime non-split links of or fewe…
The sparse group lasso optimization problem is solved using a coordinate gradient descent algorithm. The algorithm is applicable to a broad class of convex loss functions. Convergence of the algorithm is established, and the algorithm is used to investigate the performance of the multinomial sparse group lasso classifi…
Efficient method for generating adversarial examples with limited query budget.
A graph is 2-apex if it is planar after the deletion of at most two vertices. Such graphs are not intrinsically knotted, IK. We investigate the converse, does not IK imply 2-apex? We determine the simplest possible counterexample, a graph on nine vertices and 21 edges that is neither IK nor 2-apex. In the process, we s…
A low-rank tensor model simplifies multi-dimensional Markov chains.
Fewer data weight updates lead to faster convergence in machine learning models.
Paper provides first theoretical guarantees for hyperbolic space learning.
Deep neural networks excel at learning the training data, but often provide incorrect and confident predictions when evaluated on slightly different test examples. This includes distribution shifts, outliers, and adversarial examples. To address these issues, we propose Manifold Mixup, a simple regularizer that encoura…
The goal of a decision-based adversarial attack on a trained model is to generate adversarial examples based solely on observing output labels returned by the targeted model. We develop HopSkipJumpAttack, a family of algorithms based on a novel estimate of the gradient direction using binary information at the decision…
Examples are given of prime Legendrian knots in the standard contact 3-space that have arbitrarily many distinct Chekanov polynomials, refuting a conjecture of Lenny Ng. These are constructed using a new `Legendrian tangle replacement' technique. This technique is then used to show that the phenomenon of multiple Cheka…
Study shows semi-supervised learning can be more robust with fewer labeled examples.
Recent work on deep neural network pruning has shown there exist sparse subnetworks that achieve equal or improved accuracy, training time, and loss using fewer network parameters when compared to their dense counterparts. Orthogonal to pruning literature, deep neural networks are known to be susceptible to adversarial…