MaxRR efficiently unlearns models by splitting and selecting core samples.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Dual-Channel Tensor Neural Network (DC-TNN) decomposes tensor data into low-rank and sparse components for better estimation and inference.
New algorithm detects cores in graphs with community structure, improving vertex selection for better clustering.
Bayesian Pseudo Label Selection reduces overfitting in semi-supervised learning.
Active learning is a machine learning approach for reducing the data labeling effort. Given a pool of unlabeled samples, it tries to select the most useful ones to label so that a model built from them can achieve the best possible performance. This paper focuses on pool-based sequential active learning for regression …
Laplace kernel feature selection offers statistical guarantees for nonparametric models with few samples.
We develop a general approach to valid inference after model selection. At the core of our framework is a result that characterizes the distribution of a post-selection estimator conditioned on the selection event. We specialize the approach to model selection by the lasso to form valid confidence intervals for the sel…
Optimizes model training efficiency with core subset selection.
Paper simulates LR fuzzy intervals with interval-valued cores.
We present the Parallel, Forward-Backward with Pruning (PFBP) algorithm for feature selection (FS) in Big Data settings (high dimensionality and/or sample size). To tackle the challenges of Big Data FS PFBP partitions the data matrix both in terms of rows (samples, training examples) as well as columns (features). By e…
Data selection methods, such as active learning and core-set selection, are useful tools for machine learning on large datasets. However, they can be prohibitively expensive to apply in deep learning because they depend on feature representations that need to be learned. In this work, we show that we can greatly improv…
Bayesian optimization technique scaled using Vecchia approximations.
Bayesian tensor train kernel machine uses Laplace approximation for scalable GP regression.
TA-CQR predicts regression intervals with exact coverage, splitting miscoverage between endpoints.
New method improves CATE model selection with optimal regret rates.
The objective of this work is to study the applicability of various Machine Learning algorithms for prediction of some rock properties which geoscientists usually define due to special lab analysis. We demonstrate that these special properties can be predicted only basing on routine core analysis (RCA) data. To validat…
A new method uses Shapley values to select important features for classification.
Paper introduces PTL-SI for statistical inference in TL-HDR, controlling FPR.
Recent work by Brock et al. (2018) suggests that Generative Adversarial Networks (GANs) benefit disproportionately from large mini-batch sizes. Unfortunately, using large batches is slow and expensive on conventional hardware. Thus, it would be nice if we could generate batches that were effectively large though actual…
CORES2 removes noisy labels by sieving out corrupted examples.
SelMix fine-tunes pre-trained models to optimize non-decomposable objectives.
We define two non-linear operations with random (not necessarily closed) sets in Banach space: the conditional core and the conditional convex hull. While the first is sublinear, the second one is superlinear (in the reverse set inclusion ordering). Furthermore, we introduce the generalised conditional expectation of r…
Bayesian optimization improves policy search in reinforcement learning.
This study examines cores within superclusters, highlighting their transitional nature and dynamical state.
Unified notation simplifies information-theoretic concepts in machine learning.
Word2vec is a widely used algorithm for extracting low-dimensional vector representations of words. State-of-the-art algorithms including those by Mikolov et al. have been parallelized for multi-core CPU architectures, but are based on vector-vector operations with "Hogwild" updates that are memory-bandwidth intensive …
Convolutional neural networks (CNNs) have been successfully applied to many recognition and learning tasks using a universal recipe; training a deep model on a very large dataset of supervised examples. However, this approach is rather restrictive in practice since collecting a large set of labeled images is very expen…
Customer scoring models are the core of scalable direct marketing. Uplift models provide an estimate of the incremental benefit from a treatment that is used for operational decision-making. Training and monitoring of uplift models require experimental data. However, the collection of data under randomized treatment as…
In statistical connectomics, the quantitative study of brain networks, estimating the mean of a population of graphs based on a sample is a core problem. Often, this problem is especially difficult because the sample or cohort size is relatively small, sometimes even a single subject. While using the element-wise sampl…
Data collection often involves the partial measurement of a larger system. A common example arises in collecting network data: we often obtain network datasets by recording all of the interactions among a small set of core nodes, so that we end up with a measurement of the network consisting of these core nodes along w…
Novel privatization framework for high-dimensional variable selection with differential privacy.
Small deformations of marginally outer trapped surfaces (MOTS) are studied by using the stability operator introduced by Andersson-Mars-Simon. Novel formulae for the principal eigenvalue are presented. A characterization of the many marginally outer trapped tubes (MOTT) passing through a given MOTS is given, and the po…
We detail distributed algorithms for scalable, secure multiparty linear regression and feature selection at essentially the same speed as plaintext regression. While the core geometric ideas are simple, the recognition of their broad utility when combined is novel. Our scheme opens the door to efficient and secure geno…
New algorithm for contextual dueling bandits achieves nearly optimal regret.
New algorithms for model selection in linear bandits adapt to instance complexity.
MARS automatically selects tensor decomposition ranks, improving performance in neural network tasks.
A new method speeds up ALS for recommender systems by subsampling key elements.
Core-Halo solves large-scale fixed-point problems by decentralizing updates.
New algorithm FLUTE achieves uniform-PAC convergence in RL with linear approx.
New core inflation measure predicts future headline inflation.
SwISS improves scalability of Bayesian inference for large datasets.
A genome-wide association study (GWAS) correlates marker variation with trait variation in a sample of individuals. Each study subject is genotyped at a multitude of SNPs (single nucleotide polymorphisms) spanning the genome. Here we assume that subjects are unrelated and collected at random and that trait values are n…
This work improves independence tests for high-dimensional data.
We develop a framework for post model selection inference, via marginal screening, in linear regression. At the core of this framework is a result that characterizes the exact distribution of linear functions of the response , conditional on the model being selected (``condition on selection" framework). This allows…
abess efficiently solves various machine learning problems quickly.
We design and mathematically analyze sampling-based algorithms for regularized loss minimization problems that are implementable in popular computational models for large data, in which the access to the data is restricted in some way. Our main result is that if the regularizer's effect does not become negligible as th…
The implied volatility surface (IVS) is a fundamental building block in computational finance. We provide a survey of methodologies for constructing such surfaces. We also discuss various topics which can influence the successful construction of IVS in practice: arbitrage-free conditions in both strike and time, how to…
Study trade-offs between statistical and computational efficiency in variational inference.