Improves pre-trial risk assessments by making them safer without changing existing rules.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We explore the problem of learning under selective labels in the context of algorithm-assisted decision making. Selective labels is a pervasive selection bias problem that arises when historical decision making blinds us to the true outcome for certain instances. Examples of this are common in many applications, rangin…
Predictive modeling is increasingly being employed to assist human decision-makers. One purported advantage of replacing human judgment with computer models in high stakes settings-- such as sentencing, hiring, policing, college admissions, and parole decisions-- is the perceived "neutrality" of computers. It is argued…
Recidivism prediction provides decision makers with an assessment of the likelihood that a criminal defendant will reoffend that can be used in pre-trial decision-making. It can also be used for prediction of locations where crimes most occur, profiles that are more likely to commit violent crimes. While such instrumen…
New framework predicts earnings announcements using press release content, surpassing earnings surprises.
Privacy is enhanced by synthetic data release even with unlimited data.
New algorithm maintains privacy while improving model performance in selective release.
A new method for releasing AI workflows to avoid premature incorrect results.
Automates phased release strategy to balance risk and speed.
User releases data to service provider while balancing privacy and utility.
GRAND ensures node-level differential privacy for network data.
New algorithms improve privacy-preserving data release using external predictions.
We lay theoretical foundations for new database release mechanisms that allow third-parties to construct consistent estimators of population statistics, while ensuring that the privacy of each individual contributing to the database is protected. The proposed framework rests on two main ideas. First, releasing (an esti…
BioFinBERT analyzes sentiment of biotech press releases and financial text around inflection points.
Proposes minimal interventions over counterfactual explanations for algorithmic recourse.
A new method for private query release using Johnson-Lindenstrauss projection.
We introduce GraSPy, a Python library devoted to statistical inference, machine learning, and visualization of random graphs and graph populations. This package provides flexible and easy-to-use algorithms for analyzing and understanding graphs with a scikit-learn compliant API. GraSPy can be downloaded from Python Pac…
This paper provides a method for noise-calibrated inference from DP synthetic data.
The ability to analyze and forecast stratospheric weather conditions is fundamental to addressing climate change. However, our capacity to collect data in the stratosphere is limited by sparsely deployed weather balloons. We propose a framework to collect stratospheric data by releasing a contrail of tiny sensor device…
Study shows monetary policy impacts digital assets like BTC and ETH.
New algorithm corrects bias in LDP-released data for better analysis.
Releasing full data records is one of the most challenging problems in data privacy. On the one hand, many of the popular techniques such as data de-identification are problematic because of their dependence on the background knowledge of adversaries. On the other hand, rigorous methods such as the exponential mechanis…
We propose a novel computational strategy for de novo design of molecules with desired properties termed ReLeaSE (Reinforcement Learning for Structural Evolution). Based on deep and reinforcement learning approaches, ReLeaSE integrates two deep neural networks - generative and predictive - that are trained separately b…
The identification of sources of advection-diffusion transport is based usually on solving complex ill-posed inverse models against the available state- variable data records. However, if there are several sources with different locations and strengths, the data records represent mixtures rather than the separate influ…
AI and HPC help screen millions of molecules for SARS-CoV-2 treatments.
This paper documents the release of the ELKI data mining framework, version 0.7.5. ELKI is an open source (AGPLv3) data mining software written in Java. The focus of ELKI is research in algorithms, with an emphasis on unsupervised methods in cluster analysis and outlier detection. In order to achieve high performance a…
The study provides a practical strategy for pricing and hedging equity-release mortgages guarantees.
Differential privacy of Gaussian process posterior sampling
What makes a paper independently reproducible? Debates on reproducibility center around intuition or assumptions but lack empirical results. Our field focuses on releasing code, which is important, but is not sufficient for determining reproducibility. We take the first step toward a quantifiable answer by manually att…
This work addresses privacy issues in IoT data sharing by balancing information disclosure and user privacy.
Additive noise protects privacy in releasing datasets for SVM classification.
In a unified framework we study equilibrium in the presence of an insider having information on the signal of the firm value, which is naturally connected to the fundamental price of the firm related asset. The fundamental value itself is announced at a future random (stopping) time. We consider two cases. First when t…
Three new oracle-efficient algorithms for private synthetic data release.
Designing a data sharing mechanism without sacrificing too much privacy can be considered as a game between data holders and malicious attackers. This paper describes a compressive adversarial privacy framework that captures the trade-off between the data privacy and utility. We characterize the optimal data releasing …
Differential privacy is a framework for privately releasing summaries of a database. Previous work has focused mainly on methods for which the output is a finite dimensional vector, or an element of some discrete set. We develop methods for releasing functions while preserving differential privacy. Specifically, we sho…
Study private query release with public data, reducing sample sizes.
Framework purifies approximate differential privacy to pure differential privacy.
Private release of sensitive data enables fair learning.
Smart Meters (SMs) are able to share the power consumption of users with utility providers almost in real-time. These fine-grained signals carry sensitive information about users, which has raised serious concerns from the privacy viewpoint. In this paper, we focus on real-time privacy threats, i.e., potential attacker…
Twitter releases dataset to study user engagement on Home Timeline, focusing on privacy.
mvlearn simplifies multiview machine learning for non-specialists.
For a dataset of label-count pairs, an anonymized histogram is the multiset of counts. Anonymized histograms appear in various potentially sensitive contexts such as password-frequency lists, degree distribution in social networks, and estimation of symmetric properties of discrete distributions. Motivated by these app…
Paper addresses private online convex optimization with optimal algorithms in various geometries and high-dimensional bandits.
Framework for efficient statistical estimation with privacy guarantees.
Differential privacy allows quantifying privacy loss resulting from accessing sensitive personal data. Repeated accesses to underlying data incur increasing loss. Releasing data as privacy-preserving synthetic data would avoid this limitation, but would leave open the problem of designing what kind of synthetic data. W…
Locally learned synaptic failure enables complete Bayesian inference.
XAI-Bench releases synthetic datasets for evaluating feature attribution methods.
Study investigates micro-event detection on FLOSS version releases from Stack Overflow.