SADCBO optimizes contextual variables by balancing relevance and cost.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Choppy optimizes ranked list truncation using Transformer architecture.
We propose in this contribution a method for l one regularization in prototype based relevance learning vector quantization (LVQ) for sparse relevance profiles. Sparse relevance profiles in hyperspectral data analysis fade down those spectral bands which are not necessary for classification. In particular, we consider …
The relevance of data quantifies learning efficiency.
Adequate evaluation of an information retrieval system to estimate future performance is a crucial task. Area under the ROC curve (AUC) is widely used to evaluate the generalization of a retrieval system. However, the objective function optimized in many retrieval systems is the error rate and not the AUC value. This p…
We design an optimal strategy for investment in a portfolio of assets subject to a multiplicative Brownian motion. The strategy provides the maximal typical long-term growth rate of investor's capital. We determine the optimal fraction of capital that an investor should keep in risky assets as well as weights of differ…
TRAIL improves robot imitation learning by focusing on task-relevant features.
DMNL bandits optimize assortment choices balancing relevance and diversity.
Variable selection for Gaussian process models is often done using automatic relevance determination, which uses the inverse length-scale parameter of each input variable as a proxy for variable relevance. This implicitly determined relevance has several drawbacks that prevent the selection of optimal input variables i…
A framework for document classification using keywords and unlabeled data.
In a market with stochastic investment opportunities, we study an optimal consumption investment problem for an agent with recursive utility of Epstein-Zin type. Focusing on the empirically relevant specification where both risk aversion and elasticity of intertemporal substitution are in excess of one, we characterize…
New approach quantifies overfitting in high-dimensional regression.
PRI-VAE learns disentangled representations by optimizing principle-of-relevant-information.
A new algorithm removes stale observations in dynamic Bayesian optimization.
RFFNet scales kernel methods to large datasets by learning kernel relevance.
AdEMAMix optimizer improves model performance and convergence speed.
We propose a method for finding alternate features missing in the Lasso optimal solution. In ordinary Lasso problem, one global optimum is obtained and the resulting features are interpreted as task-relevant features. However, this can overlook possibly relevant features not selected by the Lasso. With the proposed met…
The problem of distributed representation learning is one in which multiple sources of information are processed separately so as to learn as much information as possible about some ground truth . We investigate this problem from information-theoretic grounds, through a generalization of Tishby's ce…
In the classic sparsity-driven problems, the fundamental L-1 penalty method has been shown to have good performance in reconstructing signals for a wide range of problems. However this performance relies on a good choice of penalty weight which is often found from empirical experiments. We propose an algorithm called t…
This monograph presents the main complexity theorems in convex optimization and their corresponding algorithms. Starting from the fundamental theory of black-box optimization, the material progresses towards recent advances in structural optimization and stochastic optimization. Our presentation of black-box optimizati…
Jointly learns feature and sample relevancies for robust sparse recovery.
E-Commerce (E-Com) search is an emerging important new application of information retrieval. Learning to Rank (LETOR) is a general effective strategy for optimizing search engines, and is thus also a key technology for E-Com search. While the use of LETOR for web search has been well studied, its use for E-Com search h…
In this work, we provide a framework linking microstructural properties of an asset to the tick value of the exchange. In particular, we bring to light a quantity, referred to as implicit spread, playing the role of spread for large tick assets, for which the effective spread is almost always equal to one tick. The rel…
We analyze different re-ranking algorithms for diversification and show that majority of them are based on maximizing submodular/modular functions from the class of parameterized concave/linear over modular functions. We study the optimality of such algorithms in terms of the `total curvature'. We also show that by adj…
Proposes UICR to improve novelty in recommendation systems without sacrificing relevance.
A new jump diffusion regime-switching model is introduced, which allows for linking jumps in asset prices with regime changes. We prove the existence and uniqueness of the solution to the risk-sensitive asset management criterion maximisation problem in this setting. We provide an ODE for the optimal value function, wh…
New framework improves worst-case generalization bounds for stochastic optimization.
Novel metrics improve machine learning models for ICU patient care.
The muti-layer information bottleneck (IB) problem, where information is propagated (or successively refined) from layer to layer, is considered. Based on information forwarded by the preceding layer, each stage of the network is required to preserve a certain level of relevance with regards to a specific hidden variab…
New meta-RL method avoids exploration-exploitation trade-off.
BONSAI optimizes parameters while respecting a default configuration, reducing unnecessary changes.
An unsupervised anomaly detection method for irregularly sampled time-series data.
Many sequential decision-making tasks require choosing at each decision step the right action out of the vast set of possibilities by extracting actionable intelligence from high-dimensional data streams. Most of the times, the high-dimensionality of actions and data makes learning of the optimal actions by traditional…
New method selects relevant dimensions for better prediction in mixtures.
Optimizes Airbnb pricing to increase revenue and reduce booking regret.
Feature selection has been proven a powerful preprocessing step for high-dimensional data analysis. However, most state-of-the-art methods tend to overlook the structural correlation information between pairwise samples, which may encapsulate useful information for refining the performance of feature selection. Moreove…
Data coarse graining improves model performance by filtering out less relevant features.
The paper debiases mini-batch approximations in deep learning for more accurate optimization and uncertainty quantification.
In this paper, we implement multi-label neural networks with optimal thresholding to identify gas species among a multi gas mixture in a cluttered environment. Using infrared absorption spectroscopy and tested on synthesized spectral datasets, our approach outperforms conventional binary relevance - partial least squar…
PAMA learns covariate importance for better matching in observational studies.
A new method optimizes recommender systems by addressing missing-not-at-random implicit feedback.
New optimal prior avoids bias in complex models with limited data.
RID framework quantifies and regularizes task-relevant knowledge in distillation.
XGBoost fails to accurately identify relevant features, while interpretable methods do.
Swift-Sarsa combines TD learning with Sarsa to control tasks robustly.
Algorithm learns diverse rankings for search engines.
Improved time complexity for parallel stochastic optimization in heterogeneous systems.
Since the events of the Arab Spring, there has been increased interest in using social media to anticipate social unrest. While efforts have been made toward automated unrest prediction, we focus on filtering the vast volume of tweets to identify tweets relevant to unrest, which can be provided to downstream users for …