New meta-optimizer learns from both point-based and population-based algorithms.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Deep neural networks (DNNs) have transformed several artificial intelligence research areas including computer vision, speech recognition, and natural language processing. However, recent studies demonstrated that DNNs are vulnerable to adversarial manipulations at testing time. Specifically, suppose we have a testing …
Although there has been substantial research in software analytics for effort estimation in traditional software projects, little work has been done for estimation in agile projects, especially estimating user stories or issues. Story points are the most common unit of measure used for estimating the effort involved in…
Novel framework for efficient Gaussian process models with monotonicity constraints.
We recall the notion of Jacobi fields, as it was extended to systems of second-order ordinary differential equations. Two points along a base integral curve are conjugate if there exists a non-trivial Jacobi field along that curve that vanishes on both points. Based on arguments that involve the eigendistributions of t…
We consider a family of variational problems on a Hilbert manifold parameterized by an open subset of a Banach manifold, and we discuss the genericity of the nondegeneracy condition for the critical points. Based on an idea of B. White, we prove an abstract genericity result that employs the infinite dimensional Sard--…
Develops accelerated fixed-point methods with delayed oracles for scientific computing.
We construct kernel, which generalizes the classical Gaussian RBF kernel to the case of incomplete data. We model the uncertainty contained in missing attributes making use of data distribution and associate every point with a conditional probability density function. This allows to embed incomplete data i…
Develops forecast hedging for improved calibration of forecasts.
New method uses random convex polytopes to measure representation quality.
Quantization based techniques are the current state-of-the-art for scaling maximum inner product search to massive databases. Traditional approaches to quantization aim to minimize the reconstruction error of the database points. Based on the observation that for a given query, the database points that have the largest…
Finite mixture models have become a popular tool for clustering. Amongst other uses, they have been applied for clustering longitudinal data and clustering high-dimensional data. In the latter case, a latent Gaussian mixture model is sometimes used. Although there has been much work on clustering using latent variables…
In this article we extend to generic -energy minimizing maps between Riemannian manifolds a regularity result which is known to hold in the case . We first show that the set of singular points of such a map can be quantitatively stratified: we classify singular points based on the number of almost-symmetries of…
An important task in machine learning and statistics is the approximation of a probability measure by an empirical measure supported on a discrete point set. Stein Points are a class of algorithms for this task, which proceed by sequentially minimising a Stein discrepancy between the empirical measure and the target an…
Adaptive selection of IPs improves online GP performance.
We develop parallel predictive entropy search (PPES), a novel algorithm for Bayesian optimization of expensive black-box objective functions. At each iteration, PPES aims to select a batch of points which will maximize the information gain about the global maximizer of the objective. Well known strategies exist for sug…
Bayesian reinforcement learning (BRL) encodes prior knowledge of the world in a model and represents uncertainty in model parameters by maintaining a probability distribution over them. This paper presents Monte Carlo BRL (MC-BRL), a simple and general approach to BRL. MC-BRL samples a priori a finite set of hypotheses…
The study formalizes temporal precision and recall for anomaly detection in sequences.
In this paper, we investigate a divide and conquer approach to Kernel Ridge Regression (KRR). Given n samples, the division step involves separating the points based on some underlying disjoint partition of the input space (possibly via clustering), and then computing a KRR estimate for each partition. The conquering s…
Proposes a new loss function for learning with noisy labels.
3D dataset for intracranial aneurysms aids deep learning applications.
This thesis uses Kantorovich-Rubinstein distance for classifying points based on their measures.
Efficient synthetic data generation improves model performance on tabular data.
We take the novel perspective to view data not as a probability distribution but rather as a current. Primarily studied in the field of geometric measure theory, -currents are continuous linear functionals acting on compactly supported smooth differential forms and can be understood as a generalized notion of orient…
Improved sparse Gaussian processes using structured scaling matrices and Power-EP framework.
In point-based sensing systems such as coordinate measuring machines (CMM) and laser ultrasonics where complete sensing is impractical due to the high sensing time and cost, adaptive sensing through a systematic exploration is vital for online inspection and anomaly quantification. Most of the existing sequential sampl…
A new method captures higher-order interactions in data clusters.
Given a solution to a linear homogeneous second order elliptic equation with Lipschitz coefficients, we introduce techniques for giving improved estimates of the critical set $\Cr(u)\equiv \{x:|\nabla u|(x)=0\}$. The results are new even for harmonic functions on $\dR^n$. Given such a , the standard {\it first o…
Efficient NTF algorithm for large sparse tensors.
Electronic records contain sequences of events, some of which take place all at once in a single visit, and others that are dispersed over multiple visits, each with a different timestamp. We postulate that fine temporal detail, e.g., whether a series of blood tests are completed at once or in rapid succession should n…
ANNs solve financial option valuation problems without numerical methods.
Unsupervised domain mapping has attracted substantial attention in recent years due to the success of models based on the cycle-consistency assumption. These models map between two domains by fooling a probabilistic discriminator, thereby matching the probability distributions of the real and generated data. Instead of…
New method finds points for approximating distributions faster.
New fairness measures account for prediction uncertainties to detect bias.
Conventional hardware-friendly quantization methods, such as fixed-point or integer, tend to perform poorly at very low word sizes as their shrinking dynamic ranges cannot adequately capture the wide data distributions commonly seen in sequence transduction models. We present AdaptivFloat, a floating-point inspired num…
DLFM models complex systems with uncertainty, outperforming traditional methods.
Optimizes risk assessment tools using mixed-integer programming.
This paper reviews metrics to assess AI model calibration accuracy.
The performance of many machine learning techniques depends on the choice of an appropriate similarity or distance measure on the input space. Similarity learning (or metric learning) aims at building such a measure from training data so that observations with the same (resp. different) label are as close (resp. far) a…
Algorithm selects variables and bandwidths for geographically weighted regression.
GCAO improves clustering of high-dimensional data by grouping low-density boundary points.
Model change points in time-series data with neural SDEs and variational autoencoders.
A new hierarchical clustering method selects representative points from sub-minimum-spanning-trees.
Missing values frequently arise in modern biomedical studies due to various reasons, including missing tests or complex profiling technologies for different omics measurements. Missing values can complicate the application of clustering algorithms, whose goals are to group points based on some similarity criterion. A c…
A new framework for mining high utility patterns in interval-based sequences.
Modern classification problems frequently present mild to severe label imbalance as well as specific requirements on classification characteristics, and require optimizing performance measures that are non-decomposable over the dataset, such as F-measure. Such measures have spurred much interest and pose specific chall…
MoReL models multi-omics data to find hidden molecular interactions.
Paper improves robustness of PINNs by smoothing and quantifying uncertainty.