Study identifies negative data externalities affecting model performance on specific groups.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study finds more flood risk strategies can improve outcomes in NYC.
Deep neural networks (DNNs) transform stimuli across multiple processing stages to produce representations that can be used to solve complex tasks, such as object recognition in images. However, a full understanding of how they achieve this remains elusive. The complexity of biological neural networks substantially exc…
One of the promising methods for the treatment of complex diseases such as cancer is combinational therapy. Due to the combinatorial complexity, machine learning models can be useful in this field, where significant improvements have recently been achieved in determination of synergistic combinations. In this study, we…
Signed pairwise interactions conflate uniqueness, redundancy, and synergy
Study how neural networks learn from non-Gaussian data models.
A scalable framework uses Langevin sampling to approximate neural network models of evolving processes.
Historical (Stressed-) Value-at-Risk ((S)VAR), and Expected Shortfall (ES), are widely used risk measures in regulatory capital and Initial Margin, i.e. funding, computations. However, whilst the definitions of VAR and ES are unambiguous, they depend on input distributions that are data-cleaning- and Data-Model-depende…
Developed Taylor series for muscle-finger system analysis.
Skilled robotic manipulation benefits from complex synergies between non-prehensile (e.g. pushing) and prehensile (e.g. grasping) actions: pushing can help rearrange cluttered objects to make space for arms and fingers; likewise, grasping can help displace objects to make pushing movements more precise and collision-fr…
Paper improves SOMs for non-Euclidean data modeling.
DL models can outperform regionalized models in hydrology by pooling diverse data.
This paper proposes a deep learning model combining CNN and Transformer for improved credit default prediction.
The problem of learning a sparse model is conceptually interpreted as the process of identifying active features/samples and then optimizing the model over them. Recently introduced safe screening allows us to identify a part of non-active features/samples. So far, safe screening has been individually studied either fo…
The abstract discusses how humans use visualizations in machine learning.
Responding to discussions on missing data models.
In our study, we demonstrate the synergy effect between convolutional neural networks and the multiplicity of SMILES. The model we propose, the so-called Convolutional Neural Fingerprint (CNF) model, reaches the accuracy of traditional descriptors such as Dragon (Mauri et al. [22]), RDKit (Landrum [18]), CDK2 (Willigha…
Unified perspective unites Bayesian optimization and active learning for efficient goal-oriented optimization.
Proposes CoDEAL for estimating heterogeneous treatment effects in panel data models.
Estimates interactions between modalities for multimodal data.
Our article considers the class of recently developed stochastic models that combine claims payments and incurred losses information into a coherent reserving methodology. In particular, we develop a family of Heirarchical Bayesian Paid-Incurred-Claims models, combining the claims reserving models of Hertig et al. (198…
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
Efficient tensor decomposition for count data models achieves near-optimal multiway analysis.
Breiman's data analysis dichotomy is outdated, offering a third approach: mechanistic models.
In Divide & Recombine (D&R), big data are divided into subsets, each analytic method is applied to subsets, and the outputs are recombined. This enables deep analysis and practical computational performance. An innovate D\&R procedure is proposed to compute likelihood functions of data-model (DM) parameters for big dat…
ecpc R-package improves high-dimensional prediction with co-data.
PDTS improves robustness in sequential decision-making.
Enhances inverse design optimization with machine learning and reduced fidelity simulations.
Understanding how different information sources together transmit information is crucial in many domains. For example, understanding the neural code requires characterizing how different neurons contribute unique, redundant, or synergistic pieces of information about sensory or behavioral variables. Williams and Beer (…
New method disentangles high-order effects in feature importance.
PoPPy is a Point Process toolbox based on PyTorch, which achieves flexible designing and efficient learning of point process models. It can be used for interpretable sequential data modeling and analysis, e.g., Granger causality analysis of multi-variate point processes, point process-based simulation and prediction of…
In this work we present a review of the state of the art of information theoretic feature selection methods. The concepts of feature relevance, redundance and complementarity (synergy) are clearly defined, as well as Markov blanket. The problem of optimal feature selection is defined. A unifying theoretical framework i…
The paper proposes a method to estimate heterogeneous treatment effects using pretraining strategies.
Commentary on Rashomon Effect complicating model selection.
Missing data is a pervasive problem in data analyses, resulting in datasets that contain censored realizations of a target distribution. Many approaches to inference on the target distribution using censored observed data, rely on missing data models represented as a factorization with respect to a directed acyclic gra…
High-dimensional data models, often with low sample size, abound in many interdisciplinary studies, genomics and large biological systems being most noteworthy. The conventional assumption of multinormality or linearity of regression may not be plausible for such models which are likely to be statistically complex due …
Leo Breiman's Rashomon Effect and Occam Dilemma are re-evaluated in the context of modern machine learning.
Proposes a new method to better understand complex system interactions.
Weak supervision challenges black-box models, suggesting fusion of modeling cultures.
TriTPP models enable faster and more flexible event data modeling.
We aim to predict and explain service failures in supply-chain networks, more precisely among last-mile pickup and delivery services to customers. We analyze a dataset of 500,000 services using (1) supervised classification with Random Forests, and (2) Association Rules. Our classifier reaches an average sensitivity of…
Combining causality, control, and reinforcement learning for system control.
R package xtdml uses DML for panel data models with fixed effects.
The Gaussian mixture model is a classic technique for clustering and data modeling that is used in numerous applications. With the rise of big data, there is a need for parameter estimation techniques that can handle streaming data and distribute the computation over several processors. While online variants of the Exp…
A faster Bayesian method for estimating spatial count data models.
New -algebra approach unifies machine learning strategies.
Robust low-rank matrix estimation is a topic of increasing interest, with promising applications in a variety of fields, from computer vision to data mining and recommender systems. Recent theoretical results establish the ability of such data models to recover the true underlying low-rank matrix when a large portion o…
New theory explains how equivariant self-supervised learning improves feature extraction.