Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

2925858771,169 · Jun 202019922001200920172026
48 results for data notions

We present a study of generalization for data-dependent hypothesis sets. We give a general learning guarantee for data-dependent hypothesis sets based on a notion of transductive Rademacher complexity. Our main result is a generalization bound for data-dependent hypothesis sets expressed in terms of a notion of hypothe…

2019-04-09abs ↗pdf ↗

We give a completely formalized definition of a notion of " general manifold ". It turns out that " gluing data " form an equivalence-partially ordered set (e-pos), which is a special instance of an ordered groupoid. We state and prove reconstruction theorems, allowing to reconstruct general manifolds and their mor-phi…

2016-05-25abs ↗pdf ↗

The paper introduces the Banzhaf value for robust data valuation in machine learning, addressing stochastic model performance.

problem Inconsistent data value rankings due to model performance noise.
method Introduces the Banzhaf value and Maximum Sample Reuse (MSR) principle for efficient estimation.
result The Banzhaf value outperforms other semivalues in robust data valuation.

Adaptive data analysis is frequently criticized for its pessimistic generalization guarantees. The source of these pessimistic bounds is a model that permits arbitrary, possibly adversarial analysts that optimally use information to bias results. While being a central issue in the field, still lacking are notions of na…

2019-01-30abs ↗pdf ↗

Statistical tests for fairness in admissions data reveal hidden patterns.

problem Simpson's paradox in admissions data hides true gender bias.
method Introduces a new statistical test based on Pearl's instrumental-variable inequalities.
result Statistical tests for fairness coincide with causal notions for the Berkeley admissions case.

New uniqueness concept for adversarial Bayes classifier.

problem Understanding adversarial Bayes classifiers in binary classification.
method Developed a new notion of uniqueness and analyzed it for a family of one-dimensional data distributions.
result Improved regularity of adversarial Bayes classifiers as perturbation radius increases.

Paper defines and proves geometric uniqueness of Einstein field equations.

problem Einstein field equations characteristic Cauchy problem
method Covariant definition of double null data, proving geometric uniqueness
result Double null data fully covariant and geometrically unique

We address the problem of verifying neural-based perception systems implemented by convolutional neural networks. We define a notion of local robustness based on affine and photometric transformations. We show the notion cannot be captured by previously employed notions of robustness. The method proposed is based on re…

2018-11-28abs ↗pdf ↗

New welfare-based fairness notions align with existing error rate balance and predictive parity.

problem Aligning fairness notions with welfare-based criteria.
method Discussing and establishing conditions for envy freeness and prejudice freeness.
result Envy freeness and prejudice freeness are equivalent to error rate balance and predictive parity.

DsDm selects data to improve model performance, avoiding handpicked notions of quality.

problem Selecting data for model training can lead to worse performance than random selection.
method Formulates dataset selection as an optimization problem, maximizing model performance.
result Selected datasets improve language model performance by 2x over baseline methods.

The adoption of automated, data-driven decision making in an ever expanding range of applications has raised concerns about its potential unfairness towards certain social groups. In this context, a number of recent studies have focused on defining, detecting, and removing unfairness from data-driven decision systems. …

2017-06-30abs ↗pdf ↗

The vector of periodic, compound returns of a typical investment portfolio is almost never a convex combination of the return vectors of the securities in the portfolio. As a result the ex post version of Harry Markowitz's "standard mean-variance portfolio selection model" does not apply to compound return data. We pro…

2011-04-28abs ↗pdf ↗

Paper investigates optimal interpolation methods in linear regression.

problem Understanding when interpolating methods generalize well in linear regression.
method Investigates optimal response-linear interpolators using functions linear in the response variable.
result Provides a closed-form expression for the optimal interpolator and shows it can be derived as the limit of gradient descent.

Generalized statistical arbitrage concepts are introduced corresponding to trading strategies which yield positive gains on average in a class of scenarios rather than almost surely. The relevant scenarios or market states are specified via an information system given by a σσ-algebra and so this notion contains classi…

2019-07-22abs ↗pdf ↗

"How much is my data worth?" is an increasingly common question posed by organizations and individuals alike. An answer to this question could allow, for instance, fairly distributing profits among multiple data contributors and determining prospective compensation when data breaches happen. In this paper, we study the…

2019-02-27abs ↗pdf ↗

The Bartnik mass is a notion of quasi-local mass which is remarkably difficult to compute. Mantoulidis and Schoen [2016] developed a novel technique to construct asymptotically flat extensions of minimal Bartnik data in such a way that the ADM mass of these extensions is well-controlled, and thus, they were able to com…

2019-03-21abs ↗pdf ↗

Sparse NMF with archetypal regularization aims to robustly represent data points.

problem Representing data points as sparse linear combinations of archetypes.
method Sparse NMF with archetypal regularization, introducing strong and weak robustness.
result Theoretical robustness guarantees hold under minimal assumptions.

With the growing interest on Network Analysis, Relational Data Mining is becoming an emphasized domain of Data Mining. This paper addresses the problem of extracting representative elements from a relational dataset. After defining the notion of degree of representativeness, computed using the Borda aggregation procedu…

2012-07-03abs ↗pdf ↗

Fairness in algorithmic decision-making processes is attracting increasing concern. When an algorithm is applied to human-related decision-making an estimator solely optimizing its predictive power can learn biases on the existing data, which motivates us the notion of fairness in machine learning. while several differ…

2018-06-13abs ↗pdf ↗

Data analytics and machine learning techniques are being rapidly adopted into the power system, including power system control as well as electricity market design. In this paper, from an adversarial machine learning point of view, we examine the vulnerability of data-driven electricity market design. More precisely, w…

2019-11-18abs ↗pdf ↗

A new model for defective media using two scales.

problem Modeling defects in media with two scales.
method Generalization of Riemann-Cartan manifolds and fibre bundle theory, constructing a first-order placement map.
result Emergent behaviors like dislocations and disclinations arise from the interaction of macroscopic and microscopic scales.

Private statistics estimation faces a bias, accuracy, and privacy trilemma.

problem Balancing privacy, accuracy, and bias in statistical estimation.
method Use differential privacy (DP) for private statistics, but clip samples to control sensitivity and add noise for privacy, introducing bias.
result No algorithm can simultaneously have low bias, low error, and low privacy loss for arbitrary distributions.

The problem of clustering is considered, for the case when each data point is a sample generated by a stationary ergodic process. We propose a very natural asymptotic notion of consistency, and show that simple consistent algorithms exist, under most general non-parametric assumptions. The notion of consistency is as f…

2010-05-05abs ↗pdf ↗

The problem of clustering is considered, for the case when each data point is a sample generated by a stationary ergodic process. We propose a very natural asymptotic notion of consistency, and show that simple consistent algorithms exist, under most general non-parametric assumptions. The notion of consistency is as f…

2010-04-29abs ↗pdf ↗

Fairness of classification and regression has received much attention recently and various, partially non-compatible, criteria have been proposed. The fairness criteria can be enforced for a given classifier or, alternatively, the data can be adapated to ensure that every classifier trained on the data will adhere to d…

2019-11-15abs ↗pdf ↗

In this follow up work to [45, 33, 32, 46] we introduce and study a notion of geodesic stability restricted to rays with prescribed singularity types. A number of notions of interest fit into this framework, in particular algebraic- and transcendental K-polystability, equivariant K-polystability, and the geodesic K-pol…

2018-12-28abs ↗pdf ↗

Clustering is an essential data mining tool that aims to discover inherent cluster structure in data. As such, the study of clusterability, which evaluates whether data possesses such structure, is an integral part of cluster analysis. Yet, despite their central role in the theory and application of clustering, current…

2016-02-22abs ↗pdf ↗

Paper tackles sample-efficient offline RL, proposing data diversity and unified algorithms.

problem Sample-efficient learning from historical data for sequential decision-making.
method Proposes data diversity and unifies three offline RL algorithm classes: VS, RO, and PS.
result Comparable sample efficiency for VS, RO, and PS algorithms under standard assumptions.

Clustering methods group a set of data points into a few coherent groups or clusters of similar data points. As an example, consider clustering pixels in an image (or video) if they belong to the same object. Different clustering methods are obtained by using different notions of similarity and different representation…

2019-11-18abs ↗pdf ↗

Study on new hyperbolicity notions for non-Kähler manifolds and their deformations.

problem Analyzing new hyperbolicity notions for non-Kähler complex manifolds.
method Introducing and analyzing two new notions of hyperbolicity for compact complex non-Kähler manifolds, and studying their behavior under smooth modifications.
result Established openness results for pp-HS hyperbolicity and pp-Kähler hyperbolicity under holomorphic deformations.

Identifying meaningful signal buried in noise is a problem of interest arising in diverse scenarios of data-driven modeling. We present here a theoretical framework for exploiting intrinsic geometry in data that resists noise corruption, and might be identifiable under severe obfuscation. Our approach is based on uncov…

2018-01-25abs ↗pdf ↗

The paper introduces group-representative clustering to ensure fair representation of different groups in clusters.

problem Ensuring fair representation of different groups in clusters.
method Developed a new clustering approach called group-representative clustering, which parallels fairness notions in classification.
result Presented approximation algorithms for group representative kk-median clustering and evaluated on real-world data.