Algorithm learns from both labeled and arbitrary test examples, giving guarantees for bounded VC dimension classes.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Generalizes Fenchel conjugation to nonlinear functions on arbitrary sets.
Algorithm learns binary function efficiently under arbitrary covariate shift.
Recently, the binary expansion testing framework was introduced to test the independence of two continuous random variables by utilizing symmetry statistics that are complete sufficient statistics for dependence. We develop a new test based on an ensemble approach that uses the sum of squared symmetry statistics and di…
Like all other knot polynomials, the superpolynomials should be defined in arbitrary representation R of the gauge group in (refined) Chern-Simons theory. However, not a single example is yet known of a superpolynomial beyond symmetric or antisymmetric representations. We consider the expansion of the superpolynomial a…
This paper presents the R package MCS which implements the Model Confidence Set (MCS) procedure recently developed by Hansen et al. (2011). The Hansen's procedure consists on a sequence of tests which permits to construct a set of 'superior' models, where the null hypothesis of Equal Predictive Ability (EPA) is not rej…
Few-shot classification is the task of predicting the category of an example from a set of few labeled examples. The number of labeled examples per category is called the number of shots (or shot number). Recent works tackle this task through meta-learning, where a meta-learner extracts information from observed tasks …
The paper develops robust tests for detecting independence in synchronous stochastic systems with finite sample guarantees.
e-LOND algorithm controls FDR in online testing with arbitrary dependencies.
We exploit a recently derived inversion scheme for arbitrary deep neural networks to develop a new semi-supervised learning framework that applies to a wide range of systems and problems. The approach outperforms current state-of-the-art methods on MNIST reaching of test set accuracy while using labeled e…
Paper presents robust confidence sequences for means with known moment bounds and arbitrary corruption.
We consider the problem of asynchronous online testing, aimed at providing control of the false discovery rate (FDR) during a continual stream of data collection and testing, where each test may be a sequential test that can start and stop at arbitrary times. This setting increasingly characterizes real-world applicati…
We present a general framework, the coupled compound Poisson factorization (CCPF), to capture the missing-data mechanism in extremely sparse data sets by coupling a hierarchical Poisson factorization with an arbitrary data-generating model. We derive a stochastic variational inference algorithm for the resulting model …
Method constructs confidence regions for linear models with arbitrary predictors.
Introduces valuative stability for polarised varieties, equivalent to K-stability.
Adversarial examples are maliciously perturbed inputs designed to mislead machine learning (ML) models at test-time. They often transfer: the same adversarial example fools more than one model. In this work, we propose novel methods for estimating the previously unknown dimensionality of the space of adversarial inputs…
The quadric ansatz solves dKP equations in arbitrary dimensions, leading to Einstein-Weyl structures.
There is a significant literature on methods for incorporating knowledge into multiple testing procedures so as to improve their power and precision. Some common forms of prior knowledge include (a) beliefs about which hypotheses are null, modeled by non-uniform prior weights; (b) differing importances of hypotheses, m…
Study predicts SGD test loss for structured features.
Study optimal ridge regularization for out-of-distribution prediction.
Proves sufficiency of countable test plans for BV functions on metric spaces.
Machine learning algorithms are known to be susceptible to data poisoning attacks, where an adversary manipulates the training data to degrade performance of the resulting classifier. In this work, we present a unifying view of randomized smoothing over arbitrary functions, and we leverage this novel characterization t…
On any surface we give an example of a metric that contains simple closed geodesics with arbitrary high Morse index. Similarly, on any 3-manifold we give an example of a metric that contains embedded minimal tori with arbitrary high Morse index. Previously no such examples were known. We also discuss whether or not suc…
The structure of a Bayesian network includes a great deal of information about the probability distribution of the data, which is uniquely identified given some general distributional assumptions. Therefore it's important to study its variability, which can be used to compare the performance of different learning algor…
Deep generative models are rapidly becoming a common tool for researchers and developers. However, as exhaustively shown for the family of discriminative models, the test-time inference of deep neural networks cannot be fully controlled and erroneous behaviors can be induced by an attacker. In the present work, we show…
This paper introduces a new framework for data efficient and versatile learning. Specifically: 1) We develop ML-PIP, a general framework for Meta-Learning approximate Probabilistic Inference for Prediction. ML-PIP extends existing probabilistic interpretations of meta-learning to cover a broad class of methods. 2) We i…
Algorithm predicts with optimal loss by abstaining from uncertain test examples.
Deep neural networks (DNNs) have transformed several artificial intelligence research areas including computer vision, speech recognition, and natural language processing. However, recent studies demonstrated that DNNs are vulnerable to adversarial manipulations at testing time. Specifically, suppose we have a testing …
In this work, we give the first algorithms for tolerant testing of nontrivial classes in the active model: estimating the distance of a target function to a hypothesis class C with respect to some arbitrary distribution D, using only a small number of label queries to a polynomial-sized pool of unlabeled examples drawn…
Extends knockoff filter for composite null hypotheses in variable selection.
We relate the existence problem of harmonic maps into to the convex geometry of . On one hand, this allows us to construct new examples of harmonic maps of degree 0 from compact surfaces of arbitrary genus into . On the other hand, we produce new example of regions that do not contain closed geodesics (…
There has been significant study on the sample complexity of testing properties of distributions over large domains. For many properties, it is known that the sample complexity can be substantially smaller than the domain size. For example, over a domain of size , distinguishing the uniform distribution from distrib…
The convolutional neural network (CNN) architecture is increasingly being applied to new domains, such as malware detection, where it is able to learn malicious behavior from raw bytes extracted from executables. These architectures reach impressive performance with no feature engineering effort involved, but their rob…
New method controls false discoveries in online testing with deadlines.
New insights on bias in multi-armed bandits under conditional sampling.
The reproducing kernel Hilbert space (RKHS) embedding of distributions offers a general and flexible framework for testing problems in arbitrary domains and has attracted considerable amount of attention in recent years. To gain insights into their operating characteristics, we study here the statistical performance of…
Real-world machine learning applications often have complex test metrics, and may have training and test data that are not identically distributed. Motivated by known connections between complex test metrics and cost-weighted learning, we propose addressing these issues by using a weighted loss function with a standard…
A new sequential test for unnormalized densities.
In the present paper, we consider a position vector of an arbitrary curve in the three-dimensional Galilean 3-space. Furthermore, we give some conditions on the curvatures of this arbitrary curve to study special curves and their Smarandache curves. Finally, in the light of this study, some related examples of these cu…
We propose a novel procedure which adds "content-addressability" to any given unconditional implicit model e.g., a generative adversarial network (GAN). The procedure allows users to control the generative process by specifying a set (arbitrary size) of desired examples based on which similar samples are generated from…
New examples show some convex-cocompact subgroups are separable.
Methods to find correlation among variables are of interest to many disciplines, including statistics, machine learning, (big) data mining and neurosciences. Parameters that measure correlation between two variables are of limited utility when used with multiple variables. In this work, I propose a simple criterion to …
Culler-Shalen theory extended to arbitrary characteristic.
Test log-likelihood comparisons can be misleading.
Model-X test detects conditional independence in streaming data.
Auckly gave two examples of irreducible integer homology spheres (one toroidal and one hyperbolic) which are not surgery on a knot in the three-sphere. Using Heegaard Floer homology, the authors and Karakurt provided infinitely many small Seifert fibered examples. In this note, we extend those results to give infinitel…
Develops a test to distinguish between standard and rough volatility.
A machine learning model that generalizes well should obtain low errors on unseen test examples. Thus, if we know how to optimally perturb training examples to account for test examples, we may achieve better generalization performance. However, obtaining such perturbation is not possible in standard machine learning f…