A new method for releasing AI workflows to avoid premature incorrect results.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work establishes always-valid risk bounds for online matrix completion.
Paper proposes a new dynamic pricing method with always-valid online statistical learning.
Within the setup of continuous-time semimartingale financial markets, we show that a multiprior Gilboa-Schmeidler minimax expected utility maximizer forms a portfolio consisting only of the riskless asset if and only if among the investor's priors there exists a probability measure under which all admissible wealth pro…
New framework predicts earnings announcements using press release content, surpassing earnings surprises.
Privacy is enhanced by synthetic data release even with unlimited data.
The paper provides bounds on the CDF of a variable under nonstationary conditions.
Deep generative models have been wildly successful at learning coherent latent representations for continuous data such as video and audio. However, generative modeling of discrete data such as arithmetic expressions and molecular structures still poses significant challenges. Crucially, state-of-the-art methods often …
New algorithm maintains privacy while improving model performance in selective release.
Automates phased release strategy to balance risk and speed.
User releases data to service provider while balancing privacy and utility.
Proof of local well-posedness for a specific boundary condition in general relativity.
GRAND ensures node-level differential privacy for network data.
New algorithms improve privacy-preserving data release using external predictions.
We lay theoretical foundations for new database release mechanisms that allow third-parties to construct consistent estimators of population statistics, while ensuring that the privacy of each individual contributing to the database is protected. The proposed framework rests on two main ideas. First, releasing (an esti…
BioFinBERT analyzes sentiment of biotech press releases and financial text around inflection points.
A new method for private query release using Johnson-Lindenstrauss projection.
The paper addresses statistical inference issues in adaptive experiments.
We introduce GraSPy, a Python library devoted to statistical inference, machine learning, and visualization of random graphs and graph populations. This package provides flexible and easy-to-use algorithms for analyzing and understanding graphs with a scikit-learn compliant API. GraSPy can be downloaded from Python Pac…
This paper provides a method for noise-calibrated inference from DP synthetic data.
The ability to analyze and forecast stratospheric weather conditions is fundamental to addressing climate change. However, our capacity to collect data in the stratosphere is limited by sparsely deployed weather balloons. We propose a framework to collect stratospheric data by releasing a contrail of tiny sensor device…
Study shows monetary policy impacts digital assets like BTC and ETH.
New algorithm corrects bias in LDP-released data for better analysis.
Releasing full data records is one of the most challenging problems in data privacy. On the one hand, many of the popular techniques such as data de-identification are problematic because of their dependence on the background knowledge of adversaries. On the other hand, rigorous methods such as the exponential mechanis…
We propose a novel computational strategy for de novo design of molecules with desired properties termed ReLeaSE (Reinforcement Learning for Structural Evolution). Based on deep and reinforcement learning approaches, ReLeaSE integrates two deep neural networks - generative and predictive - that are trained separately b…
The identification of sources of advection-diffusion transport is based usually on solving complex ill-posed inverse models against the available state- variable data records. However, if there are several sources with different locations and strengths, the data records represent mixtures rather than the separate influ…
AI and HPC help screen millions of molecules for SARS-CoV-2 treatments.
This paper documents the release of the ELKI data mining framework, version 0.7.5. ELKI is an open source (AGPLv3) data mining software written in Java. The focus of ELKI is research in algorithms, with an emphasis on unsupervised methods in cluster analysis and outlier detection. In order to achieve high performance a…
We propose an alternative framework to existing setups for controlling false alarms when multiple A/B tests are run over time. This setup arises in many practical applications, e.g. when pharmaceutical companies test new treatment options against control pills for different diseases, or when internet companies test the…
The study provides a practical strategy for pricing and hedging equity-release mortgages guarantees.
Differential privacy of Gaussian process posterior sampling
What makes a paper independently reproducible? Debates on reproducibility center around intuition or assumptions but lack empirical results. Our field focuses on releasing code, which is important, but is not sufficient for determining reproducibility. We take the first step toward a quantifiable answer by manually att…
Given only positive (P) and unlabeled (U) data, PU learning can train a binary classifier without any negative data. It has two building blocks: PU class-prior estimation (CPE) and PU classification; the latter has been well studied while the former has received less attention. Hitherto, the distributional-assumption-f…
This work addresses privacy issues in IoT data sharing by balancing information disclosure and user privacy.
In a unified framework we study equilibrium in the presence of an insider having information on the signal of the firm value, which is naturally connected to the fundamental price of the firm related asset. The fundamental value itself is announced at a future random (stopping) time. We consider two cases. First when t…
Three new oracle-efficient algorithms for private synthetic data release.
Designing a data sharing mechanism without sacrificing too much privacy can be considered as a game between data holders and malicious attackers. This paper describes a compressive adversarial privacy framework that captures the trade-off between the data privacy and utility. We characterize the optimal data releasing …
Differential privacy is a framework for privately releasing summaries of a database. Previous work has focused mainly on methods for which the output is a finite dimensional vector, or an element of some discrete set. We develop methods for releasing functions while preserving differential privacy. Specifically, we sho…
Study private query release with public data, reducing sample sizes.
Framework purifies approximate differential privacy to pure differential privacy.
Private release of sensitive data enables fair learning.
Smart Meters (SMs) are able to share the power consumption of users with utility providers almost in real-time. These fine-grained signals carry sensitive information about users, which has raised serious concerns from the privacy viewpoint. In this paper, we focus on real-time privacy threats, i.e., potential attacker…
Twitter releases dataset to study user engagement on Home Timeline, focusing on privacy.
We derive concentration inequalities for differentially private median and mean estimators building on the "Propose, Test, Release" (PTR) mechanism introduced by Dwork and Lei (2009). We introduce a new general version of the PTR mechanism that allows us to derive high probability error bounds for differentially privat…
mvlearn simplifies multiview machine learning for non-specialists.
For a dataset of label-count pairs, an anonymized histogram is the multiset of counts. Anonymized histograms appear in various potentially sensitive contexts such as password-frequency lists, degree distribution in social networks, and estimation of symmetric properties of discrete distributions. Motivated by these app…
Paper addresses private online convex optimization with optimal algorithms in various geometries and high-dimensional bandits.
Framework for efficient statistical estimation with privacy guarantees.