Improves model classification accuracy in black-box settings.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper we address cardinality estimation problem which is an important subproblem in query optimization. Query optimization is a part of every relational DBMS responsible for finding the best way of the execution for the given query. These ways are called plans. The execution time of different plans may differ b…
New method reduces regret in budgeted learning problems.
Discrete integration in a high dimensional space of n variables poses fundamental challenges. The WISH algorithm reduces the intractable discrete integration problem into n optimization queries subject to randomized constraints, obtaining a constant approximation guarantee. The optimization queries are expensive, which…
Proposes a new query autocompletion method that maximizes retrieval performance.
Game theory helps machine learn better from adversarial queries.
Deep neural networks have recently achieved tremendous success in image classification. Recent studies have however shown that they are easily misled into incorrect classification decisions by adversarial examples. Adversaries can even craft attacks by querying the model in black-box settings, where no information abou…
This paper improves retrieval for LLMs in financial document Q&A.
The problem of active diagnosis arises in several applications such as disease diagnosis, and fault diagnosis in computer networks, where the goal is to rapidly identify the binary states of a set of objects (e.g., faulty or working) by sequentially selecting, and observing, (noisy) responses to binary valued queries. …
A classical problem in causal inference is that of matching, where treatment units need to be matched to control units based on covariate information. In this work, we propose a method that computes high quality almost-exact matches for high-dimensional categorical datasets. This method, called FLAME (Fast Large-scale …
Federated learning is a distributed form of machine learning where both the training data and model training are decentralized. In this paper, we use federated learning in a commercial, global-scale setting to train, evaluate and deploy a model to improve virtual keyboard search suggestion quality without direct access…
LITE models improve query-document relevance with learnable late interactions.
We address the two fundamental problems of spatial field reconstruction and sensor selection in heterogeneous sensor networks: (i) how to efficiently perform spatial field reconstruction based on measurements obtained simultaneously from networks with both high and low quality sensors; and (ii) how to perform query bas…
Recent studies have shown that adversarial examples in state-of-the-art image classifiers trained by deep neural networks (DNN) can be easily generated when the target model is transparent to an attacker, known as the white-box setting. However, when attacking a deployed machine learning service, one can only acquire t…
Active learning selects high-quality examples for text-to-SQL systems.
Pairwise "same-cluster" queries are one of the most widely used forms of supervision in semi-supervised clustering. However, it is impractical to ask human oracles to answer every query correctly. In this paper, we study the influence of allowing "not-sure" answers from a weak oracle and propose an effective algorithm …
Establishes upper bounds on generalization error in active learning.
Optimizes query routing to LLMs under cost and resource constraints.
A new attack for probabilistic classifiers adapts to noise levels.
Unified framework for deferring queries to top-k experts, improving accuracy-cost trade-offs.
Machine learning models have been found to be susceptible to adversarial examples that are often indistinguishable from the original inputs. These adversarial examples are created by applying adversarial perturbations to input samples, which would cause them to be misclassified by the target models. Attacks that search…
Noisy labeled data is more a norm than a rarity for self-generated content that is continuously published on the web and social media. Due to privacy concerns and governmental regulations, such a data stream can only be stored and used for learning purposes in a limited duration. To overcome the noise in this on-line s…
BMBO-DARN optimizes expensive functions with varying fidelities.
Dynamic classifier selection systems aim to select a group of classifiers that is most adequate for a specific query pattern. This is done by defining a region around the query pattern and analyzing the competence of the classifiers in this region. However, the regions are often surrounded by noise which can difficult …
Meta-Router optimizes LLM selection using gold-standard and preference-based data.
The article presents a method that improves the quality of classification of objects described by a combination of known and unknown features. The method is based on modernized Informational Neurobayesian Approach with consideration of unknown features. The proposed method was developed and trained on 1500 text queries…
Bayesian method helps decision-makers find preferred solutions in multi-objective optimization.
Estimates neural architecture performance speedily.
Scarcity of training data for task-oriented dialogue systems is a well known problem that is usually tackled with costly and time-consuming manual data annotation. An alternative solution is to rely on automatic text generation which, although less accurate than human supervision, has the advantage of being cheap and f…
Improves retrieval accuracy for hierarchical documents, especially for distant matches.
The web contains a vast corpus of HTML tables. They can be used to provide direct answers to many web queries. We focus on answering two classes of queries with those tables: those seeking lists of entities (e.g., `cities in california') and those seeking superlative entities (e.g., `largest city in california'). The m…
Query2box embeds complex queries as boxes to handle logical operations in large KGs.
Constraint-based clustering algorithms exploit background knowledge to construct clusterings that are aligned with the interests of a particular user. This background knowledge is often obtained by allowing the clustering system to pose pairwise queries to the user: should these two elements be in the same cluster or n…
Collaborative filtering (CF) allows the preferences of multiple users to be pooled to make recommendations regarding unseen products. We consider in this paper the problem of online and interactive CF: given the current ratings associated with a user, what queries (new ratings) would most improve the quality of the rec…
Study learns linear utility functions from comparisons, showing learnability gaps between passive and active learning.
Study uses LLMs to improve financial forecasting by integrating textual and numerical data.
VMGP extends Gaussian processes for Bayesian meta-learning, improving uncertainty prediction.
Improved algorithm for selecting a hypothesis locally privately with fewer queries.
We study the problem of estimating a set of linear queries with respect to some unknown distribution over a domain based on a sensitive data set of individuals under the constraint of local differential privacy. This problem subsumes a wide range of estimation tasks, e.g., distrib…
In recent years, deep metric learning has achieved promising results in learning high dimensional semantic feature embeddings where the spatial relationships of the feature vectors match the visual similarities of the images. Similarity search for images is performed by determining the vectors with the smallest distanc…
NP-Attack reduces query counts for black-box adversarial attacks.
Solves TOD systems' query annotation problem without explicit annotations.
Proposes a max-utility arm selection strategy for reducing cumulative regret in sequential query recommendations.
A new method for selecting high quality itemsets from a large collection.
C-IP improves LLMs' query selection for interactive tasks by estimating uncertainty robustly.
Study on recovering sparse linear classifiers from mixed binary responses.
Study efficient sequential evaluation of large language models using historical data.
Study exact community recovery in noisy SBM with limited queries.