Efficiently selects nearest neighbors for labeling to speed up active learning.
problem Intractable active learning and search for large-scale unlabeled data.
method Restricts candidate pool to nearest neighbors of labeled set.
result Achieved similar performance to global approach but reduced computational cost by up to 3 orders of magnitude.
SQWA improves low-precision DNNs with model averaging and quantization.
problem Designing good generalization DNNs with quantized weights.
method Floating-point model training, direct quantization, multiple low-precision models, weight averaging, re-quantization, fine-tuning, loss visualization.
result State-of-the-art results for 2-bit QDNNs on CIFAR-100 and ImageNet datasets.
Low precision operations can provide scalability, memory savings, portability, and energy efficiency. This paper proposes SWALP, an approach to low precision training that averages low-precision SGD iterates with a modified learning rate schedule. SWALP is easy to implement and can match the performance of full-precisi…
The paper uses deep neural networks to estimate and infer ATE without needing to know the dimension of the data.
problem Estimating and inferring the average treatment effect (ATE) in complex data settings.
method The paper uses deep neural networks to estimate the mean regression function and then calculates the ATE. It establishes consistency and asymptotic normality of the estimators.
result The deep neural network estimates of ATE are consistent and asymptotically normal, providing dimension-free rates.
Bayesian method for estimating ATE with robustness to model misspecification.
problem Estimating average treatment effects under unconfoundedness.
method Double robust Bayesian inference using adjusted prior and posterior distributions.
result Bayesian credible sets form asymptotically exact confidence intervals.
Introduces robust and decomposable AP for image retrieval.
problem Challenges in training deep neural networks with AP.
method Differentiable rank approximation and loss function design.
result ROADMAP outperforms AP approximation methods and deep models.
Study on stochastic approximation with Polyak-Ruppert averaging for linear systems.
problem Understanding the asymptotic and non-asymptotic properties of stochastic approximation procedures.
method Detailed analysis of linear stochastic approximation with Polyak-Ruppert averaging, focusing on asymptotic and non-asymptotic properties.
result Proves CLT and non-asymptotic concentration inequality for averaged iterates, providing refined understanding of linear stochastic approximation.
Outlier detection methods have become increasingly relevant in recent years due to increased security concerns and because of its vast application to different fields. Recently, Pauwels and Lasserre (2016) noticed that the sublevel sets of the inverse Christoffel function accurately depict the shape of a cloud of data …
The paper examines the consistency of item embeddings in recommendation systems.
problem The relevance of averaging item embeddings for user or concept representation.
method Proposes an expected precision score to measure consistency and analyzes it theoretically and empirically.
result Real-world averages are less consistent for recommendation compared to theoretical assumptions.
We consider a discrete-time approximation of paths of an Ornstein--Uhlenbeck process as a mean for estimation of a price of European call option in the model of financial market with stochastic volatility. The Euler--Maruyama approximation scheme is implemented. We determine the estimates for the option price for prede…
RF models implicitly regularize kernel methods as feature count increases.
problem Understanding implicit regularization in RF models.
method Random matrix theory applied to Gaussian RF models and KRR.
result The average RF predictor is close to a KRR predictor with an effective ridge.
In causal inference, a variety of causal effect estimands have been studied, including the sample, uncensored, target, conditional, optimal subpopulation, and optimal weighted average treatment effects. Ad-hoc methods have been developed for each estimand based on inverse probability weighting (IPW) and on outcome regr…
New bounds for average graph distance using curvature and centrality.
problem Finding bounds for average graph distance.
method Using weighted average Ollivier curvature with edge betweenness centrality.
result Equality in bounds achieved for specific reflective graphs.
This paper provides new insight into maximizing F1 scores in the context of binary classification and also in the context of multilabel classification. The harmonic mean of precision and recall, F1 score is widely used to measure the success of a binary classifier when one class is rare. Micro average, macro average, a…
Measures collectivity in financial covariances and correlations to reveal trends and precursors.
problem Capturing collective motion in financial markets to predict trends and precursors.
method Measures collectivity using the largest eigenvalue and average sector collectivity.
result Identifies collective signals around major financial events and captures trends in covariances and correlations.
Study precise asymptotics of noncompact Type-IIb solutions to mean curvature flow.
problem Understanding the behavior of noncompact Type-IIb solutions to mean curvature flow as time approaches infinity.
method Constructed rotationally symmetric solutions with specific asymptotic behavior and analyzed their properties.
result The highest curvature concentrates at the tip of the hypersurface and blows up at the Type-IIb rate (2t+1)(γ−1)/2. This paper optimizes Bayesian estimation for log-concave models using Langevin Monte-Carlo.
problem Optimizing Bayesian estimators for log-concave models with Langevin Monte-Carlo.
method Quantitative statistical bounds and numerical approximation of Gibbs measures.
result Established optimal numerical strategy and its cost for Bayesian posterior mean approximation.
We present a method for training multi-label, massively multi-class image classification models, that is faster and more accurate than supervision via a sigmoid cross-entropy loss (logistic regression). Our method consists in embedding high-dimensional sparse labels onto a lower-dimensional dense sphere of unit-normed …
The topic modeling discovers the latent topic probability of the given text documents. To generate the more meaningful topic that better represents the given document, we proposed a new feature extraction technique which can be used in the data preprocessing stage. The method consists of three steps. First, it generate…
Improved multi-task averaging reduces mean squared error in high-dimensional data.
problem Joint estimation of multiple distributions using independent data sets.
method Exploits similarities between tasks by shrinking naive estimators towards local averages.
result The method provides a significant reduction in mean squared error, especially in high-dimensional spaces.
Optimal trend-following strategy uses simple EMA, avoiding complex cherry-picked signals.
problem Cherry-picking signals for trend-following strategies.
method Simple EMA for trend capture, avoiding complex indicators.
result Simple EMA is optimal for capturing trend, complex indicators are risky.
Group averaging boosts model accuracy without training cost.
problem Challenging training of equivariant models in physics.
method Group averaging at test time, improving accuracy.
result Improves model accuracy by up to 37% in continuous dynamics.
This paper compares average-K and top-K classification methods under ambiguity.
problem Choosing a single label in ambiguous cases leads to low precision.
method Formally characterizes ambiguity profiles and compares average-K and top-K classification methods.
result Average-K can achieve lower error rates than top-K in some ambiguous cases.
The study calculates the average genus of rational knots and links.
problem Finding the average genus of rational knots and links.
method Enumerating and calculating the number of rational knots and links with a given crossing number.
result A precise formula for the average minimal genus of rational knots and links.
Paper proposes a mean-field gradient descent for zero-sum games, proving convergence to Nash equilibrium.
problem Finding mixed Nash equilibria in zero-sum games with multiple players.
method Mean-field gradient descent dynamics with time-averaging, incorporating exponentially discounted gradients.
result Exponential convergence rate to mixed Nash equilibrium with respect to total variation metric.
The paper develops optimal confidence regions for categorical data.
problem Constructing tight confidence regions for categorical data.
method Develops new theory for minimum average volume confidence regions.
result Shows optimality of the regions for categorical data and its implications for machine learning.
Hyperbolic embeddings offer excellent quality with few dimensions when embedding hierarchical data structures like synonym or type hierarchies. Given a tree, we give a combinatorial construction that embeds the tree in hyperbolic space with arbitrarily low distortion without using optimization. On WordNet, our combinat…
A new method for averaging data on manifolds is proposed, offering simplicity and efficiency.
problem The difficulty of computing Fréchet means on manifolds, especially Stiefel and Grassmann.
method Proposed RL-barycenters, simpler arithmetic means projected onto the manifold.
result RL-barycenters yield simple yet effective means on Stiefel and Grassmann manifolds.
HAVER improves error bounds for estimating the largest mean in machine learning tasks.
problem Estimating the largest mean among multiple distributions.
method Proposes HAVER, a novel algorithm for maximum mean estimation.
result HAVER achieves better error bounds than the oracle in many cases.
In a rotationally symmetric space $\oM$ around an axis A (whose precise definition includes all real space forms), we consider a domain G limited by two equidistant hypersurfaces orthogonal to A. Let $M \subset \oM$ be a revolution hypersurface generated by a graph over A, with boundary in ∂G and orthogonal…
Estimates scalar curvature on moduli space as genus grows.
problem Estimating scalar curvature on moduli space.
method Analyzing Weil-Petersson metric for large genus.
result Precise estimate of average scalar curvature up to 1/g2. In many machine learning applications, crowdsourcing has become the primary means for label collection. In this paper, we study the optimal error rate for aggregating labels provided by a set of non-expert workers. Under the classic Dawid-Skene model, we establish matching upper and lower bounds with an exact exponent …
MCRapper efficiently computes patterns in data using Monte-Carlo Rademacher Averages.
problem Finding statistically significant patterns in data with limited samples.
method Monte-Carlo Empirical Rademacher Averages (MCERA) for poset families.
result MCRapper provides upper bounds to the discrepancy of functions, enabling efficient pattern mining.
The paper analyzes real-time methods to detect rapidly varying liquidity in markets.
problem Increased trade execution price uncertainty due to rapid price variations by high-frequency traders.
method A four-state Markov switching model to identify volatile liquidity states.
result The model can generate a signal to delay orders, reducing price volatility for market participants.
We study the problem of large scale, multi-label visual recognition with a large number of possible classes. We propose a method for augmenting a trained neural network classifier with auxiliary capacity in a manner designed to significantly improve upon an already well-performing model, while minimally impacting its c…
The paper introduces new portfolio rules beyond mean-variance, addressing asymmetry and uncertainty.
problem Optimizing portfolios with asymmetric returns and uncertainty in expected returns.
method Derives allocation rules for asymmetric Laplace distributed returns and random normal expected returns. Addresses singular covariance matrices and uncertainty in returns.
result Optimal worst-case scenario solution provides a convex alternative to risk parity, improving portfolio stability.
The optimal ranking score between precision and recall is rarely F1 and can be found using specific methods.
problem Finding a meaningful and optimal compromise between precision and recall scores.
method Established a shortest path between precision- and recall-induced rankings, framed the problem as an optimization problem, and provided theoretical tools to find the optimal β.
result F1 and its skew-insensitive version are not optimal tradeoffs between precision and recall scores.
With ever-increasing computational demand for deep learning, it is critical to investigate the implications of the numeric representation and precision of DNN model weights and activations on computational efficiency. In this work, we explore unconventional narrow-precision floating-point representations as it relates …
We study consistency properties of machine learning methods based on minimizing convex surrogates. We extend the recent framework of Osokin et al. (2017) for the quantitative analysis of consistency properties to the case of inconsistent surrogates. Our key technical contribution consists in a new lower bound on the ca…
Neural network model improves leaf spectral reflectance prediction for grapevines.
problem Inaccurate modeling of grapevine leaf spectral reflectance from traits.
method Multi-head attention neural network trained on grapevine-specific data.
result Model achieved high accuracy (R^2=0.84, NRMSE=1.52%) and outperformed PROSPECT-PRO.
Paper introduces stability in model averaging and proposes a L2-penalty method.
problem Theoretical properties of model averaging from stability perspective.
method Introduces stability, defines asymptotic empirical risk minimizer, and proposes L2-penalty model averaging method.
result Proposed L2-penalty method ensures stability and consistency under reasonable conditions.
A new meta-algorithm for estimating the conditional average treatment effects is proposed in the paper. The main idea underlying the algorithm is to consider a new dataset consisting of feature vectors produced by means of concatenation of examples from control and treatment groups, which are close to each other. Outco…
A framework estimates multiple precision matrices with shared structures.
problem Estimating multiple precision matrices with shared structures.
method Penalized likelihood framework with iterative algorithm alternating between convex and clustering problems.
result The method outperforms competitors and performs similarly to methods using prior information.
New algorithm for solving minimax problems over distributions converges to Nash equilibrium.
problem Solving minimax problems over probability distributions.
method Symmetric Mean-field Langevin Dynamics (MFL-AG and MFL-ABR) with weighted averaging and best response dynamics.
result Converges to mixed Nash equilibrium with average-iterate and last-iterate convergence.
This paper presents our system details and results of participation in the RDoC Tasks of BioNLP-OST 2019. Research Domain Criteria (RDoC) construct is a multi-dimensional and broad framework to describe mental health disorders by combining knowledge from genomics to behaviour. Non-availability of RDoC labelled dataset …
LSTM outperforms traditional models in forecasting international migration.
problem Precise forecasting of international migration for policymaking.
method Replaced a gravity linear model with an LSTM approach using Google Trends data.
result LSTM approach combined with Google Trends data outperforms existing models.
On-line portfolio selection has attracted increasing interests in machine learning and AI communities recently. Empirical evidences show that stock's high and low prices are temporary and stock price relatives are likely to follow the mean reversion phenomenon. While the existing mean reversion strategies are shown to …
Study optimal portfolio selection with Recovery Average Value at Risk, showing better control over liabilities.
problem Optimizing portfolios with a new risk measure under known or uncertain distributions.
method Existence results for mean-risk optimal portfolios under different distributional assumptions.
result Portfolio selection under Recovery Average Value at Risk provides better control over liabilities.