Study optimizes data collection from biased, costly sources to minimize risk.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We give complete algorithms and source code for constructing statistical risk models, including methods for fixing the number of risk factors. One such method is based on eRank (effective rank) and yields results similar to (and further validates) the method set forth in an earlier paper by one of us. We also give a co…
Clarinet uses complementary labels to train classifiers with less source data.
A new method reduces energy consumption in machine learning by using multiple, less costly data sources.
Adaptive source selection for positive transfer in linear models improves target dataset performance.
LLMs help less-resourced researchers access costly data.
Deep learning has produced state-of-the-art results for a variety of tasks. While such approaches for supervised learning have performed well, they assume that training and testing data are drawn from the same distribution, which may not always be the case. As a complement to this challenge, single-source unsupervised …
New method uses neural networks to identify sources from limited data in complex systems.
This paper studies the problem of stance detection which aims to predict the perspective (or stance) of a given document with respect to a given claim. Stance detection is a major component of automated fact checking. As annotating stances in different domains is a tedious and costly task, automatic methods based on ma…
Collecting large training datasets, annotated with high-quality labels, is costly and time-consuming. This paper proposes a novel framework for training deep convolutional neural networks from noisy labeled datasets that can be obtained cheaply. The problem is formulated using an undirected graphical model that represe…
Transfer learning assumes classifiers of similar tasks share certain parameter structures. Unfortunately, modern classifiers uses sophisticated feature representations with huge parameter spaces which lead to costly transfer. Under the impression that changes from one classifier to another should be ``simple'', an effi…
Develops a framework for cost-efficient Bayesian optimization with constraints.
SYNC generates synthetic data from aggregated sources using Gaussian copulas.
Transfer learning, in which a network is trained on one task and re-purposed on another, is often used to produce neural network classifiers when data is scarce or full-scale training is too costly. When the goal is to produce a model that is not only accurate but also adversarially robust, data scarcity and computatio…
New algorithms reduce costly feature collection in bandits.
Model infers mineral locations from geospatial data, improving predictions with auxiliary data.
A new compressed sensing system speeds up PSTE detection.
ConfEviSurrogate improves surrogate model accuracy and uncertainty quantification.
Large Sinkhorn couplings improve flow models in data generation tasks.
Modern machine learning systems such as image classifiers rely heavily on large scale data sets for training. Such data sets are costly to create, thus in practice a small number of freely available, open source data sets are widely used. We suggest that examining the geo-diversity of open data sets is critical before …
Surveying how to use unlabeled data in federated learning.
DP models misspecify LF dependencies, leading to significant performance errors.
Social media sources can provide crucial information in crisis situations, but discovering relevant messages is not trivial. Methods have so far focused on universal detection models for all kinds of crises or for certain crisis types (e.g. floods). Event-specific models could implement a more focused search area, but …
Study uses three sources to evaluate language models fairly.
Regulations impose idiosyncratic capital and funding costs for holding derivatives. Capital requirements are costly because derivatives desks are risky businesses; funding is costly in part because regulations increase the minimum funding tenor. Idiosyncratic costs mean no single measure makes derivatives martingales f…
Paper proposes efficient algorithms for bandit problems with costly sampling.
Although software analytics has experienced rapid growth as a research area, it has not yet reached its full potential for wide industrial adoption. Most of the existing work in software analytics still relies heavily on costly manual feature engineering processes, and they mainly address the traditional classification…
In most real-world settings such as recommender systems, finance, and healthcare, collecting useful information is costly and requires an active choice on the part of the decision maker. The decision-maker needs to learn simultaneously what observations to make and what actions to take. This paper incorporates the info…
Prediction in a small-sized sample with a large number of covariates, the "small n, large p" problem, is challenging. This setting is encountered in multiple applications, such as precision medicine, where obtaining additional samples can be extremely costly or even impossible, and extensive research effort has recentl…
Bayesian optimization tackles mixed discrete-continuous problems with Gaussian processes.
Many post-disaster and -conflict regions do not have sufficient data on their transportation infrastructure assets, hindering both mobility and reconstruction. In particular, as the number of aging and deteriorating bridges increase, it is necessary to quantify their load characteristics in order to inform maintenance …
JPEG2000 (j2k) is a highly popular format for image and video compression.With the rapidly growing applications of cloud based image classification, most existing j2k-compatible schemes would stream compressed color images from the source before reconstruction at the processing center as inputs to deep CNNs. We propose…
Novel method reduces costly model evaluations in inference problems.
Acquiring ground truth labels for unlabelled data can be a costly procedure, since it often requires manual labour that is error-prone. Consequently, the available amount of labelled data is increasingly reduced due to the limitations of manual data labelling. It is possible to increase the amount of labelled data samp…
This work analyzes and optimizes memory and compute costs of learned optimizers.
New method for reliability analysis using multi-fidelity models.
Deep learning speeds up real-time emission monitoring.
Supervised machine learning methods usually require a large set of labeled examples for model training. However, in many real applications, there are plentiful unlabeled data but limited labeled data; and the acquisition of labels is costly. Active learning (AL) reduces the labeling cost by iteratively selecting the mo…
New feature selection methods improve uplift modeling accuracy.
MAD framework learns operators from physics-embedded data efficiently.
In2Core selects a coreset for efficient LLM fine-tuning with reduced data.
LLMs simulate financial markets, revealing consistent trading strategies and market dynamics.
Survey of Monte Carlo methods for noisy, costly densities in reinforcement learning and ABC.
Theoretical understanding of deep learning is one of the most important tasks facing the statistics and machine learning communities. While deep neural networks (DNNs) originated as engineering methods and models of biological networks in neuroscience and psychology, they have quickly become a centerpiece of the machin…
Efficiently tests two distributions with few label queries.
Novel approach for estimating conditional expectations using Bayesian quadrature.
Optimal investment strategy with expert opinions in uncertain conditions.
Study predicts doubling of U.S. maize insurance claims due to climate change.