Enhances index selection for databases with task-specific inductive biases.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Tool to estimate research impact for low-resource institutions.
We study the distribution of fluctuations over a time scale (i.e., the returns) of the S&P 500 index by analyzing three distinct databases. Database (i) contains approximately 1 million records sampled at 1 min intervals for the 13-year period 1984-1996, database (ii) contains 8686 daily records for the 35-year pe…
Study compares CDS databases and finds discrepancies due to various factors.
Content based image retrieval, a technique which uses visual contents of image to search images from large scale image databases according to users' interests. This paper provides a comprehensive survey on recent technology used in the area of content based face image retrieval. Nowadays digital devices and photo shari…
CwA optimizes search performance by jointly learning a balanced database partition and a neural probing function.
New method accurately reconstructs Russell 3000 index, revealing crowded portfolios.
Geometric Brownian motion simulates stock prices for Brazilian small caps index.
In order to study the phenomenon in detail that income distribution follows Pareto law, we analyze the database of high income companies in Japan. We find a quantitative relation between the average capital of the companies and the Pareto index. The larger the average capital becomes, the smaller the Pareto index becom…
LIFT uses demonstrations to train reinforcement learning controllers for data management tasks.
Amortizes MIPS by training neural networks to predict optimal keys.
Google Trends data improves economic forecasts of private consumption.
Aligns databases with Gaussian features using MAP estimation and thresholding.
Gene expression programming predicts compression index of fine-grained soils efficiently.
We select the stocks traded in the New York Stock Exchange and we form a statistical ensemble of daily stock returns for each of the trading days of our database from the stock price time series. We study the ensemble return distribution for each trading day and we find that the symmetry properties of the ensem…
Deep neural network algorithms are difficult to analyze because they lack structure allowing to understand the properties of underlying transforms and invariants. Multiscale hierarchical convolutional networks are structured deep convolutional networks where layers are indexed by progressively higher dimensional attrib…
Study tackles database variability in medical data using ensemble models and CNNs.
Study connects database alignment and planted matching using Gaussian features.
Employing profits data of Japanese firms in 2003--2005, we kinematically exhibit the static log-normal distribution in the middle scale region. In the derivation, a Non-Gibrat's law under the detailed balance is adopted together with following two approximations. Firstly, the probability density function of profits gro…
This work applies SQL to deep learning, leveraging database techniques.
Graph database outperforms in filtering ESG stocks efficiently.
PyODDS is a Python system for outlier detection in databases.
Face recognition system trained with noisy labels.
Taxonomies of cryptocurrencies and comparisons with fiat money and databases.
Bayesian entity resolution merges together multiple, noisy databases and returns the minimal collection of unique individuals represented, together with their true, latent record values. Bayesian methods allow flexible generative models that share power across databases as well as principled quantification of uncertain…
We present a database of parliamentary debates that contains the complete record of parliamentary speeches from Dáil Éireann, the lower house and principal chamber of the Irish parliament, from 1919 to 2013. In addition, the database contains background information on all TDs (Teachta Dála, members of parliament), such…
Paper improves text-to-SQL translation by encoding schema relations with self-attention.
Graph Neural Networks improve machine learning on relational databases.
Paper presents a new Wi-Fi RSS and geomagnetic field database for indoor localization and trajectory estimation.
New deep learning method validated across multiple sleep staging databases.
EERN uses deep learning for relational databases, outperforming other methods.
In this paper we deal with the offline handwriting text recognition (HTR) problem with reduced training datasets. Recent HTR solutions based on artificial neural networks exhibit remarkable solutions in referenced databases. These deep learning neural networks are composed of both convolutional (CNN) and long short-ter…
Extends Aff-Wild database for affect recognition in real-world settings.
New vector quantization method reduces relevance of parallel components in database points.
We lay theoretical foundations for new database release mechanisms that allow third-parties to construct consistent estimators of population statistics, while ensuring that the privacy of each individual contributing to the database is protected. The proposed framework rests on two main ideas. First, releasing (an esti…
The multiple fundamental frequency detection problem and the source separation problem from a single-channel signal containing multiple oscillatory components and a nonstationary noise are both challenging tasks. To extract the fetal electrocardiogram (ECG) from a single-lead maternal abdominal ECG, we face both challe…
Topology-based information retrieval improves query accuracy.
A statistical algorithm for categorizing different types of matches and fraud in image databases is presented. The approach is based on a generative model of a graph representing images and connections between pairs of identities, trained using properties of a matching algorithm between images.
A new algorithm uses bandits to diversify database activity monitoring.
The paper studies causal effects of multiple treatments in healthcare databases with rare outcomes.
Study examines scientific research on Bitcoin across various disciplines.
The problem of searching for experts in a given academic field is hugely important in both industry and academia. We study exactly this issue with respect to a database of authors and their publications. The idea is to use Latent Semantic Indexing (LSI) and Latent Dirichlet Allocation (LDA) to perform topic modelling i…
PyODDS automates outlier detection for new data sources.
New method embeds DNA sequences for faster, more informative gene comparison.
In this paper we study predictive pattern mining problems where the goal is to construct a predictive model based on a subset of predictive patterns in the database. Our main contribution is to introduce a novel method called safe pattern pruning (SPP) for a class of predictive pattern mining problems. The SPP method a…
We propose a method to identify all the nodes that are relevant to compute all the conditional probability distributions for a given set of nodes. Our method is simple, effcient, consistent, and does not require learning a Bayesian network first. Therefore, our method can be applied to high-dimensional databases, e.g. …
We introduce multiscale invariant dictionaries to estimate quantum chemical energies of organic molecules, from training databases. Molecular energies are invariant to isometric atomic displacements, and are Lipschitz continuous to molecular deformations. Similarly to density functional theory (DFT), the molecule is re…
Hashing has been widely used for large-scale approximate nearest neighbor search because of its storage and search efficiency. Recent work has found that deep supervised hashing can significantly outperform non-deep supervised hashing in many applications. However, most existing deep supervised hashing methods adopt a …