This work redefines data-centric AI by unifying categorical and cochain notions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We introduce a new notion of the stability of computations, which holds under post-processing and adaptive composition. We show that the notion is both necessary and sufficient to ensure generalization in the face of adaptivity, for any computations that respond to bounded-sensitivity linear queries while providing acc…
We present a study of generalization for data-dependent hypothesis sets. We give a general learning guarantee for data-dependent hypothesis sets based on a notion of transductive Rademacher complexity. Our main result is a generalization bound for data-dependent hypothesis sets expressed in terms of a notion of hypothe…
We give a completely formalized definition of a notion of " general manifold ". It turns out that " gluing data " form an equivalence-partially ordered set (e-pos), which is a special instance of an ordered groupoid. We state and prove reconstruction theorems, allowing to reconstruct general manifolds and their mor-phi…
The paper introduces the Banzhaf value for robust data valuation in machine learning, addressing stochastic model performance.
Adaptive data analysis is frequently criticized for its pessimistic generalization guarantees. The source of these pessimistic bounds is a model that permits arbitrary, possibly adversarial analysts that optimally use information to bias results. While being a central issue in the field, still lacking are notions of na…
Statistical tests for fairness in admissions data reveal hidden patterns.
New method prunes classification trees for biased data.
Proposes individual fairness for clustering, making data points prefer their own cluster.
New uniqueness concept for adversarial Bayes classifier.
Paper defines and proves geometric uniqueness of Einstein field equations.
The notion of drift refers to the phenomenon that the distribution, which is underlying the observed data, changes over time. Albeit many attempts were made to deal with drift, formal notions of drift are application-dependent and formulated in various degrees of abstraction and mathematical coherence. In this contribu…
We address the problem of verifying neural-based perception systems implemented by convolutional neural networks. We define a notion of local robustness based on affine and photometric transformations. We show the notion cannot be captured by previously employed notions of robustness. The method proposed is based on re…
New distances defined between space-times, proving some definite.
New welfare-based fairness notions align with existing error rate balance and predictive parity.
DsDm selects data to improve model performance, avoiding handpicked notions of quality.
The adoption of automated, data-driven decision making in an ever expanding range of applications has raised concerns about its potential unfairness towards certain social groups. In this context, a number of recent studies have focused on defining, detecting, and removing unfairness from data-driven decision systems. …
A recent trend of fair machine learning is to define fairness as causality-based notions which concern the causal connection between protected attributes and decisions. However, one common challenge of all causality-based fairness notions is identifiability, i.e., whether they can be uniquely measured from observationa…
The vector of periodic, compound returns of a typical investment portfolio is almost never a convex combination of the return vectors of the securities in the portfolio. As a result the ex post version of Harry Markowitz's "standard mean-variance portfolio selection model" does not apply to compound return data. We pro…
Paper investigates optimal interpolation methods in linear regression.
Generalized statistical arbitrage concepts are introduced corresponding to trading strategies which yield positive gains on average in a class of scenarios rather than almost surely. The relevant scenarios or market states are specified via an information system given by a -algebra and so this notion contains classi…
"How much is my data worth?" is an increasingly common question posed by organizations and individuals alike. An answer to this question could allow, for instance, fairly distributing profits among multiple data contributors and determining prospective compensation when data breaches happen. In this paper, we study the…
In this paper, we introduce the notion of motif closure and describe higher-order ranking and link prediction methods based on the notion of closing higher-order network motifs. The methods are fast and efficient for real-time ranking and link prediction-based applications such as web search, online advertising, and re…
The Bartnik mass is a notion of quasi-local mass which is remarkably difficult to compute. Mantoulidis and Schoen [2016] developed a novel technique to construct asymptotically flat extensions of minimal Bartnik data in such a way that the ADM mass of these extensions is well-controlled, and thus, they were able to com…
Sparse NMF with archetypal regularization aims to robustly represent data points.
With the growing interest on Network Analysis, Relational Data Mining is becoming an emphasized domain of Data Mining. This paper addresses the problem of extracting representative elements from a relational dataset. After defining the notion of degree of representativeness, computed using the Borda aggregation procedu…
This research unifies concepts of fading memory in RNNs.
Rank concepts help explain deep learning's effectiveness.
Fairness in algorithmic decision-making processes is attracting increasing concern. When an algorithm is applied to human-related decision-making an estimator solely optimizing its predictive power can learn biases on the existing data, which motivates us the notion of fairness in machine learning. while several differ…
Data analytics and machine learning techniques are being rapidly adopted into the power system, including power system control as well as electricity market design. In this paper, from an adversarial machine learning point of view, we examine the vulnerability of data-driven electricity market design. More precisely, w…
Equal experience fairness in recommender systems reduces bias.
The main approach to defining equivalence among acyclic directed causal graphical models is based on the conditional independence relationships in the distributions that the causal models can generate, in terms of the Markov equivalence. However, it is known that when cycles are allowed in the causal structure, conditi…
A new model for defective media using two scales.
Private statistics estimation faces a bias, accuracy, and privacy trilemma.
The problem of clustering is considered, for the case when each data point is a sample generated by a stationary ergodic process. We propose a very natural asymptotic notion of consistency, and show that simple consistent algorithms exist, under most general non-parametric assumptions. The notion of consistency is as f…
The problem of clustering is considered, for the case when each data point is a sample generated by a stationary ergodic process. We propose a very natural asymptotic notion of consistency, and show that simple consistent algorithms exist, under most general non-parametric assumptions. The notion of consistency is as f…
Fairness of classification and regression has received much attention recently and various, partially non-compatible, criteria have been proposed. The fairness criteria can be enforced for a given classifier or, alternatively, the data can be adapated to ensure that every classifier trained on the data will adhere to d…
Develops efficient projections for multivariate probability measures.
In this follow up work to [45, 33, 32, 46] we introduce and study a notion of geodesic stability restricted to rays with prescribed singularity types. A number of notions of interest fit into this framework, in particular algebraic- and transcendental K-polystability, equivariant K-polystability, and the geodesic K-pol…
Clustering is an essential data mining tool that aims to discover inherent cluster structure in data. As such, the study of clusterability, which evaluates whether data possesses such structure, is an integral part of cluster analysis. Yet, despite their central role in the theory and application of clustering, current…
In this paper, we provide a method to learn the directed structure of a Bayesian network using data. The data is accessed by making conditional probability queries to a black-box model. We introduce a notion of simplicity of representation of conditional probability tables for the nodes in the Bayesian network, that we…
Paper tackles sample-efficient offline RL, proposing data diversity and unified algorithms.
Clustering methods group a set of data points into a few coherent groups or clusters of similar data points. As an example, consider clustering pixels in an image (or video) if they belong to the same object. Different clustering methods are obtained by using different notions of similarity and different representation…
Study on new hyperbolicity notions for non-Kähler manifolds and their deformations.
The paper proposes Tier Balancing for dynamic fairness in decision-making.
Identifying meaningful signal buried in noise is a problem of interest arising in diverse scenarios of data-driven modeling. We present here a theoretical framework for exploiting intrinsic geometry in data that resists noise corruption, and might be identifiable under severe obfuscation. Our approach is based on uncov…
Survey various symmetry notions for toric varieties.
The paper introduces group-representative clustering to ensure fair representation of different groups in clusters.