Method predicts disease outbreaks using search logs, overcoming instability.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Inference in log-linear models scales linearly in the size of output space in the worst-case. This is often a bottleneck in natural language processing and computer vision tasks when the output space is feasibly enumerable but very large. We propose a method to perform inference in log-linear models with sublinear amor…
Optimizing an interactive system against a predefined online metric is particularly challenging, when the metric is computed from user feedback such as clicks and payments. The key challenge is the counterfactual nature: in the case of Web search, any change to a component of the search engine may result in a different…
MICO uses mutual information co-training to improve selective search efficiency.
New algorithms verify and search causal graphs with minimal interventions.
New algorithm reconstructs sparse networks in subquadratic time.
AI agents improve forecast combination but require transparency.
AI agents improve forecast combination in empirical economics.
Optimistic search speeds up change point detection in large datasets.
Learning policies on data synthesized by models can in principle quench the thirst of reinforcement learning algorithms for large amounts of real experience, which is often costly to acquire. However, simulating plausible experience de novo is a hard problem for many complex environments, often resulting in biases for …
GAMA is a user-friendly AutoML system for machine learning pipeline optimization.
New algorithm tackles adversarial RL without horizon constraints.
We present a simple transformation of the formulation of the log-periodic power law formula of the Johansen-Ledoit-Sornette model of financial bubbles that reduces it to a function of only three nonlinear parameters. The transformation significantly decreases the complexity of the fitting procedure and improves its sta…
New algorithm achieves near-optimal performance in dueling bandit problem.
Previous work has shown that popular trending events are important external factors which pose significant influence on user search behavior and also provided a way to computationally model this influence. However, their problem formulation was based on the strong assumption that each event poses its influence independ…
Proposes FMS for more efficient neural network hyperparameter optimization.
Paper tackles sparse recovery with shuffled labels, establishing statistical and computational limits.
Improved DP KDE with better privacy and efficiency.
NATS-Bench benchmarks NAS algorithms for architecture topology and size.
New algorithm optimizes AUC in binary classification and changepoint detection.
In sponsored search, keyword recommendations help advertisers to achieve much better performance within limited budget. Many works have been done to mine numerous candidate keywords from search logs or landing pages. However, the strategy to select from given candidates remains to be improved. The existing relevance-ba…
We present the first sublinear memory sketch that can be queried to find the nearest neighbors in a dataset. Our online sketching algorithm compresses an N element dataset to a sketch of size in time, where . This sketch can correctly report the nearest neighbors of any …
The paper tackles feature cross search for linear models, providing approximation algorithms and structural results.
In the paper, we focus on complexity of C5.0 algorithm for constructing decision tree classifier that is the models for the classification problem from machine learning. In classical case the decision tree is constructed in running time, where is a number of classes, is the size of a traini…
This paper investigates the problem of determining a binary-valued function through a sequence of strategically selected queries. The focus is an algorithm called Generalized Binary Search (GBS). GBS is a well-known greedy algorithm for determining a binary-valued function through a sequence of strategically selected q…
We find a nonlinear dependence between an indicator of the degree of multiscaling of log-price time series of a stock and the average correlation of the stock with respect to the other stocks traded in the same market. This result is a robust stylized fact holding for different financial markets. We investigate this re…
A new method selects optimal temperature for Bayesian Deep Learning.
Mean-field variational inference is a method for approximate Bayesian posterior inference. It approximates a full posterior distribution with a factorized set of distributions by maximizing a lower bound on the marginal likelihood. This requires the ability to integrate a sum of terms in the log joint likelihood using …
Study robust learning of Lipschitz functions under corrupted binary signals.
We consider the problem of learning the structure of undirected graphical models with bounded treewidth, within the maximum likelihood framework. This is an NP-hard problem and most approaches consider local search techniques. In this paper, we pose it as a combinatorial optimization problem, which is then relaxed to a…
Estimates Schrödinger potentials with minimal sample size.
Active learning improves neutron spectroscopy experiments by automating measurement selection.
New methods solve MI problems with locally Lipschitz operators, improving solution efficiency.
Financial market prediction on the basis of online sentiment tracking has drawn a lot of attention recently. However, most results in this emerging domain rely on a unique, particular combination of data sets and sentiment tracking tools. This makes it difficult to disambiguate measurement and instrument effects from f…
Greedy MI maximization method outperforms existing approaches in nonlinear models.
Traditional approaches to ranking in web search follow the paradigm of rank-by-score: a learned function gives each query-URL combination an absolute score and URLs are ranked according to this score. This paradigm ensures that if the score of one URL is better than another then one will always be ranked higher than th…
Paper proposes a method to aggregate customer engagement data for better ranking of e-commerce results.
Research on nearest-neighbor methods tends to focus somewhat dichotomously either on the statistical or the computational aspects -- either on, say, Bayes consistency and rates of convergence or on techniques for speeding up the proximity search. This paper aims at bridging these realms: to reap the advantages of fast …
It is generally accepted that many time series of practical interest exhibit strong dependence, i.e., long memory. For such series, the sample autocorrelations decay slowly and log-log periodogram plots indicate a straight-line relationship. This necessitates a class of models for describing such behavior. A popular cl…
Modern statistical applications involving large data sets have focused attention on statistical methodologies which are both efficient computationally and able to deal with the screening of large numbers of different candidate models. Here we consider computationally efficient variational Bayes approaches to inference …
PCTS optimizes noisy, delayed, multi-fidelity feedbacks in black-box optimization.
New estimator uses clustering to improve off-policy evaluation accuracy.
With the growth of user-generated content, we observe the constant rise of the number of companies, such as search engines, content aggregators, etc., that operate with tremendous amounts of web content not being the services hosting it. Thus, aiming to locate the most important content and promote it to the users, the…
The paper solves a portfolio selection problem in incomplete markets by balancing utility and risk.
A new method detects changes in multivariate data using random forests.
Symbolic regression finds simple formulas for implied volatility.
Improved genetic programming by optimizing mutation operators for continuous program search.
Neural networks are part of many contemporary NLP systems, yet their empirical successes come at the price of vulnerability to adversarial attacks. Previous work has used adversarial training and data augmentation to partially mitigate such brittleness, but these are unlikely to find worst-case adversaries due to the c…