Modern computing and communication technologies can make data collection procedures very efficient. However, our ability to analyze large data sets and/or to extract information out from them is hard-pressed to keep up with our capacities for data collection. Among these huge data sets, some of them are not collected f…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We study a logistic model-based active learning procedure for binary classification problems, in which we adopt a batch subject selection strategy with a modified sequential experimental design method. Moreover, accompanying the proposed subject selection scheme, we simultaneously conduct a greedy variable selection pr…
We investigate active learning by pairwise similarity over the leaves of trees originating from hierarchical clustering procedures. In the realizable setting, we provide a full characterization of the number of queries needed to achieve perfect reconstruction of the tree cut. In the non-realizable setting, we rely on k…
We consider an active learning setting where the algorithm has access to a large pool of unlabeled data and a small pool of labeled data. In each iteration, the algorithm chooses few unlabeled data points and obtains their labels from an oracle. In this paper, we consider a probabilistic querying procedure to choose th…
Sparse Canonical Correlation Analysis (CCA) has received considerable attention in high-dimensional data analysis to study the relationship between two sets of random variables. However, there has been remarkably little theoretical statistical foundation on sparse CCA in high-dimensional settings despite active methodo…
Consider a set of latent factors whose observable effect of activation is caught on a measure space that appears as a grid of bits tacking value in . This paper intend to deliver a theoretical and practical answer to the question: Given that we have access to a perfect indicator of the activation of latent f…
We propose an active set selection framework for Gaussian process classification for cases when the dataset is large enough to render its inference prohibitive. Our scheme consists of a two step alternating procedure of active set update rules and hyperparameter optimization based upon marginal likelihood maximization.…
Extends active subspace analysis to infinite dimensions.
Method identifies cardiac ectopic activity sites from 12-lead ECG.
This work presents an explicit-implicit procedure to compute a model predictive control (MPC) law with guarantees on recursive feasibility and asymptotic stability. The approach combines an offline-trained fully-connected neural network with an online primal active set solver. The neural network provides a control inpu…
We study the localization of a cluster of activated vertices in a graph, from adaptively designed compressive measurements. We propose a hierarchical partitioning of the graph that groups the activated vertices into few partitions, so that a top-down sensing procedure can identify these partitions, and hence the activa…
Procedure removes training data dependency from deep networks, improving generalization.
Hidden Markov Models detect hand gestures from wearable sEMG signals.
Efficient algorithm for mobile health provides timely physical activity suggestions.
Optimizes identifying top-k items from comparisons with minimal comparisons.
Leveraging the wealth of unlabeled data produced in recent years provides great potential for improving supervised models. When the cost of acquiring labels is high, probabilistic active learning methods can be used to greedily select the most informative data points to be labeled. However, for many large-scale problem…
This paper reduces labeling costs for meta-learning in wireless systems.
New framework improves fairness in small data settings.
Echo state networks with random weights can approximate any continuous system.
Paper develops exact convex optimization for neural networks with polynomial activations.
New binary AA methods improve on existing techniques.
Proposes an efficient algorithm for mHealth that makes real-time physical activity suggestions.
Observing prices of European put and call options, we calibrate exponential Lévy models nonparametrically. We discuss the efficient implementation of the spectral estimation procedures for Lévy models of finite jump activity as well as for self-decomposable Lévy models. Based on finite sample variances, confidence inte…
Commercial activity trackers are set to become an essential tool in health research, due to increasing availability in the general population. The corresponding vast amounts of mostly unlabeled data pose a challenge to statistical modeling approaches. To investigate the feasibility of deep learning approaches for unsup…
Study challenges the Gaussian pre-activations assumption in neural networks.
Active learning method for neural population dynamics using optogenetics.
We propose a mathematical procedure for finding informed trader activities in European-style options and their underlying asset. The regression model (9) with moving average component was written. Being added to it ARMA-process for log-price differences of underlying asset, the generalized model is written as Vector AR…
A new method optimizes spatial sampling for level set estimation in one dimension.
Many iterative procedures in stochastic optimization exhibit a transient phase followed by a stationary phase. During the transient phase the procedure converges towards a region of interest, and during the stationary phase the procedure oscillates in that region, commonly around a single point. In this paper, we devel…
New algorithms reduce label collection for online prediction with expert advice.
We present safe active incremental feature selection~(SAIF) to scale up the computation of LASSO solutions. SAIF does not require a solution from a heavier penalty parameter as in sequential screening or updating the full model for each iteration as in dynamic screening. Different from these existing screening methods,…
Tempered sigmoids improve deep learning privacy.
Unified study of active learning for deep neural networks.
We present two Bayesian procedures to infer the interactions and external currents in an assembly of stochastic integrate-and-fire neurons from the recording of their spiking activity. The first procedure is based on the exact calculation of the most likely time courses of the neuron membrane potentials conditioned by …
A significantly faster algorithm is presented for the original kNN mode seeking procedure. It has the advantages over the well-known mean shift algorithm that it is feasible in high-dimensional vector spaces and results in uniquely, well defined modes. Moreover, without any additional computational effort it may yield …
New model handles missing data effectively in autoregressive models.
RAAL optimizes black box function optimization with multifidelity models.
How can we find a general way to choose the most suitable samples for training a classifier? Even with very limited prior information? Active learning, which can be regarded as an iterative optimization procedure, plays a key role to construct a refined training set to improve the classification performance in a variet…
Committee neural network models improve accuracy and enable active learning for interatomic potentials.
Study iterated ERM in active learning, deriving test error bounds.
Coop-FTPL algorithm minimizes network regret in semi-bandit settings.
Wasserstein active regression improves estimation precision.
Classification algorithms aim to predict an unknown label (e.g., a quality class) for a new instance (e.g., a product). Therefore, training samples (instances and labels) are used to deduct classification hypotheses. Often, it is relatively easy to capture instances but the acquisition of the corresponding labels remai…
Adaptive quadrature improves Bayesian inference through active learning.
Unified perspective unites Bayesian optimization and active learning for efficient goal-oriented optimization.
Layout hotpot detection is one of the main steps in modern VLSI design. A typical hotspot detection flow is extremely time consuming due to the computationally expensive mask optimization and lithographic simulation. Recent researches try to facilitate the procedure with a reduced flow including feature extraction, tra…
New algorithms for decision trees with noisy outcomes improve learning efficiency.
IUPM monitors machine learning models under gradual shifts using optimal transport and active labeling.