This paper deals with prediction of anopheles number, the main vector of malaria risk, using environmental and climate variables. The variables selection is based on an automatic machine learning method using regression trees, and random forests combined with stratified two levels cross validation. The minimum threshol…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We consider machine learning in a comparison-based setting where we are given a set of points in a metric space, but we have no access to the actual distances between the points. Instead, we can only ask an oracle whether the distance between two points and is smaller than the distance between the points an…
New framework uses MABs for better WLAN performance.
Nowadays, the major challenge in machine learning is the Big Data challenge. The big data problems due to large number of data points or large number of features in each data point, or both, the training of models have become very slow. The training time has two major components: Time to access the data and time to pro…
In2Core selects a coreset for efficient LLM fine-tuning with reduced data.
SBBO optimizes complex spaces using sampling-based models.
A popular method for selecting the number of clusters is based on stability arguments: one chooses the number of clusters such that the corresponding clustering results are "most stable". In recent years, a series of papers has analyzed the behavior of this method from a theoretical point of view. However, the results …
Framework simplifies AI access for all.
The presence of data corruption in user-generated streaming data, such as social media, motivates a new fundamental problem that learns reliable regression coefficient when features are not accessible entirely at one time. Until now, several important challenges still cannot be handled concurrently: 1) corrupted data e…
This paper focuses on the problem of determining as large a region as possible where a function exceeds a given threshold with high probability. We assume that we only have access to a noise-corrupted version of the function and that function evaluations are costly. To select the next query point, we propose maximizing…
Machine learning (ML) is probably the first and foremost used technique to deal with the size and complexity of the new generation of data. In this paper, we analyze one of the means to increase the performances of ML algorithms which is exploiting data locality. Data locality and access patterns are often at the heart…
Theorem converse to Jordan's curve theorem says that {\it if a compact set has two complementary domains in , from each of which it is at every point accessible, it is a simple closed curve}. We show that the requirement of this theorem that {\it all} points of were accessible from {\it both} complementa…
New method shows supervised learning can mimic unsupervised learning effectively.
New method for selecting data points in deep learning models.
Method finds compatible features for subsets of data.
Proposes a new method to selectively access privileged information in reinforcement learning.
The paper proposes a new model for predicting and analyzing economic variables.
Spain uses DEA to select international markets for exports.
Paper speeds up IoT device detection and data decoding.
This paper tackles belief-state selection in simulators with latent states.
Let be a polynomial of degree with a Cremer point and no repelling or parabolic periodic bi-accessible points. We show that there are two types of such Julia sets . The \emph{red dwarf} are nowhere connected im kleinen and such that the intersection of all impressions of external angles is a cont…
A new method for dynamic feature selection outperforms existing approaches.
Study of tree automorphisms via arc-stabilizers.
Inference-Time Scaling can be extended to domains prone to systematic failure using intrinsic statistics.
AMES framework selects optimal embedding space for latent graph inference.
Role mining tackles the problem of finding a role-based access control (RBAC) configuration, given an access-control matrix assigning users to access permissions as input. Most role mining approaches work by constructing a large set of candidate roles and use a greedy selection strategy to iteratively pick a small subs…
Deep generative networks can simulate from a complex target distribution, by minimizing a loss with respect to samples from that distribution. However, often we do not have direct access to our target distribution - our data may be subject to sample selection bias, or may be from a different but related distribution. W…
Bayesian Optimisation (BO) refers to a class of methods for global optimisation of a function which is only accessible via point evaluations. It is typically used in settings where is expensive to evaluate. A common use case for BO in machine learning is model selection, where it is not possible to analytically…
This paper enhances ML algorithms by improving data locality and reducing redundancy.
NUBO simplifies Bayesian optimization for researchers.
Protocol assesses better model among two opaque models with minimal interaction.
PINNACLE optimizes point selection for PINNs, improving accuracy.
For many interesting tasks, such as medical diagnosis and web page classification, a learner only has access to some positively labeled examples and many unlabeled examples. Learning from this type of data requires making assumptions about the true distribution of the classes and/or the mechanism that was used to selec…
A new RL framework optimizes drug-like molecules synthetically.
A homological selection theorem for C-spaces, as well as, a finite-dimensional homological selection theorem is established. We apply the finite-dimensional homological selection theorem to obtain fixed-point theorems for usco homologically UV^n set-valued maps.
D3M debiases models by selectively removing problematic examples.
Selects points from Jordan domains on Riemannian surfaces.
In this paper we solve the following problem in the affirmative: Let be a continuum in the plane $\complex$ and suppose that $h:Z\times [0,1]\to\complex$ is an isotopy starting at the identity. Can be extended to an isotopy of the plane? We will provide a new characterization of an accessible point in a planar …
Enhances reinforcement learning with partial state information.
New study shows data-oblivious attacks can outperform data-aware ones.
We present a sparse knowledge gradient (SpKG) algorithm for adaptively selecting the targeted regions within a large RNA molecule to identify which regions are most amenable to interactions with other molecules. Experimentally, such regions can be inferred from fluorescence measurements obtained by binding a complement…
New method speeds up distributed linear regression.
In this paper we propose a framework for automated forecasting of energy-related time series using open access data from European Network of Transmission System Operators for Electricity (ENTSO-E). The framework provides forecasts for various European countries using publicly available historical data only. Our solutio…
Method estimates causal effects from incremental data, overcoming missing data challenges.
We present a novel technique based on deep learning and set theory which yields exceptional classification and prediction results. Having access to a sufficiently large amount of labelled training data, our methodology is capable of predicting the labels of the test data almost always even if the training data is entir…
In this paper we study a family of variance reduction methods with randomized batch size---at each step, the algorithm first randomly chooses the batch size and then selects a batch of samples to conduct a variance-reduced stochastic update. We give the linear convergence rate for this framework for composite functions…
Method selects significant spatial covariates in noisy data.
PEAKS selects key training examples incrementally based on prediction error and kernel similarity.