Bayesian optimisation tackles high-dimensional categorical and mixed search spaces.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Quantum-assisted VAE improves similarity search in high-dimensional datasets.
Local PBO methods improve preferential BO in high-dimensional problems.
Optimistic search speeds up change point detection in large datasets.
LA-MCTS learns search space partition for black-box optimization using Monte Carlo Tree Search.
ICQ improves high-dimensional similarity search without sacrificing precision.
Derives a method to optimize high-dimensional functions on low-dimensional manifolds.
Novel method for high-dimensional BO using CMA to define local regions.
We develop estimation for potentially high-dimensional additive structural equation models. A key component of our approach is to decouple order search among the variables from feature or edge selection in a directed acyclic graph encoding the causal structure. We show that the former can be done with nonregularized (r…
A new line search rule improves support recovery in high-dimensional data.
The neural architecture search (NAS) algorithm with reinforcement learning can be a powerful and novel framework for the automatic discovering process of neural architectures. However, its application is restricted by noncontinuous and high-dimensional search spaces, which result in difficulty in optimization. To resol…
Direct contextual policy search methods learn to improve policy parameters and simultaneously generalize these parameters to different context or task variables. However, learning from high-dimensional context variables, such as camera images, is still a prominent problem in many real-world tasks. A naive application o…
BioHash improves similarity search performance using sparse high-dimensional hash codes.
MORBO improves multi-objective BO for high-dimensional problems.
SOLAR improves search efficiency and accuracy with sparse, orthogonal embeddings.
A new method optimizes Bayesian optimization in high dimensions by focusing on low-dimensional subspaces.
New method speeds up k-means clustering for large k by improving nearest-neighbor search.
New method reduces high-dimensional data to key features.
In this paper, we consider the problem of classification of high dimensional queries to high dimensional classes where and are discrete alphabets and the probabilistic model that relates data to the classes is known. This problem has applications …
Python package reduces hubness in high-dimensional data.
We propose a new class of data-independent locality-sensitive hashing (LSH) algorithms based on the fruit fly olfactory circuit. The fundamental difference of this approach is that, instead of assigning hashes as dense points in a low dimensional space, hashes are assigned in a high dimensional space, which enhances th…
We present a novel view of nonlinear manifold learning using derivative-free optimization techniques. Specifically, we propose an extension of the classical multi-dimensional scaling (MDS) method, where instead of performing gradient descent, we sample and evaluate possible "moves" in a sphere of fixed radius for each …
In quadruped gait learning, policy search methods that scale high dimensional continuous action spaces are commonly used. In most approaches, it is necessary to introduce prior knowledge on the gaits to limit the highly non-convex search space of the policies. In this work, we propose a new approach to encode the symme…
New algorithms improve contextual search in the presence of adversarial corruptions.
Databases are widespread, yet extracting relevant data can be difficult. Without substantial domain knowledge, multivariate search queries often return sparse or uninformative results. This paper introduces an approach for searching structured data based on probabilistic programming and nonparametric Bayes. Users speci…
This paper presents the R package gRapHD for efficient selection of high-dimensional undirected graphical models. The package provides tools for selecting trees, forests and decomposable models minimizing information criteria such as AIC or BIC, and for displaying the independence graphs of the models. It has also some…
The most data-efficient algorithms for reinforcement learning in robotics are model-based policy search algorithms, which alternate between learning a dynamical model of the robot and optimizing a policy to maximize the expected return given the model and its uncertainties. Among the few proposed approaches, the recent…
Exact risk and learning rate curves derived for adaptive SGD on high-dimensional problems.
CAGES optimizes expensive RL problems by efficiently learning gradients from multiple sources.
Proposes a Monte-Carlo method for sparse signal reconstruction.
Unsupervised space partitioning improves ANNS performance without pre-processing.
A new scheme reduces global search cost by a square root factor.
Machine learning, specifically LSTM, models quantum experiments efficiently.
Bayesian methods improve drug discovery experiment design.
A new framework scales active search for large datasets.
Bayesian optimisation with graph kernels improves neural architecture search and provides interpretability.
A new method for generating counterfactual explanations in high-dimensional datasets.
The scientific method relies on the iterated processes of inference and inquiry. The inference phase consists of selecting the most probable models based on the available data; whereas the inquiry phase consists of using what is known about the models to select the most relevant experiment. Optimizing inquiry involves …
Many emerging use cases of data mining and machine learning operate on large datasets with data from heterogeneous sources, specifically with both sparse and dense components. For example, dense deep neural network embedding vectors are often used in conjunction with sparse textual features to provide high dimensional …
Subspace clustering aims to find groups of similar objects (clusters) that exist in lower dimensional subspaces from a high dimensional dataset. It has a wide range of applications, such as analysing high dimensional sensor data or DNA sequences. However, existing algorithms have limitations in finding clusters in non-…
A new framework improves graph construction for semi-supervised learning.
KPCA-BO improves BO for high-dimensional optimization problems by learning a non-linear sub-manifold.
We develop theoretical foundations of Resonator Networks, a new type of recurrent neural network introduced in Frady et al. (2020) to solve a high-dimensional vector factorization problem arising in Vector Symbolic Architectures. Given a composite vector formed by the Hadamard product between a discrete set of high-dim…
We study the problem of treatment effect estimation in randomized experiments with high-dimensional covariate information, and show that essentially any risk-consistent regression adjustment can be used to obtain efficient estimates of the average treatment effect. Our results considerably extend the range of settings …
Unified framework connects EI and information-theoretic acquisition functions.
Consider observation data, comprised of n observation vectors with values on a set of attributes. This gives us n points in attribute space. Having data structured as a tree, implied by having our observations embedded in an ultrametric topology, offers great advantage for proximity searching. If we have preprocessed d…
Quantum algorithms improve perceptron learning efficiency.
We consider high-dimensional binary classification by sparse logistic regression. We propose a model/feature selection procedure based on penalized maximum likelihood with a complexity penalty on the model size and derive the non-asymptotic bounds for the resulting misclassification excess risk. The bounds can be reduc…