This paper quantifies privacy loss in exploratory data analysis.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The increasing availability of large but noisy data sets with a large number of heterogeneous variables leads to the increasing interest in the automation of common tasks for data analysis. The most time-consuming part of this process is the Exploratory Data Analysis, crucial for better domain understanding, data clean…
Study uses big data to analyze quantum invariants.
Making sense of a dataset in an automatic and unsupervised fashion is a challenging problem in statistics and AI. Classical approaches for {exploratory data analysis} are usually not flexible enough to deal with the uncertainty inherent to real-world data: they are often restricted to fixed latent interaction models an…
We propose in this paper an exploratory analysis algorithm for functional data. The method partitions a set of functions into clusters and represents each cluster by a simple prototype (e.g., piecewise constant). The total number of segments in the prototypes, , is chosen by the user and optimally distributed am…
New method explains high-dimensional sphere data with latent factors.
Study finds key investing characteristics for success in equity markets.
ACA identifies and explains anomalies in data.
This study analyzes data science vocabulary changes over 13 years.
Archetypal analysis helps understand binary data sets.
This paper proposes a novel profile likelihood method for estimating the covariance parameters in exploratory factor analysis of high-dimensional Gaussian datasets with fewer observations than number of variables. An implicitly restarted Lanczos algorithm and a limited-memory quasi-Newton method are implemented to deve…
A new method clusters mixed-type data tables effectively.
Action chunking and data exploration improve behavior cloning in robotics.
Paper introduces probabilistic methods to approximate archetypal analysis, reducing complexity.
This describes a statistical technique called "tonsuring" for exploratory data analysis in finance. Instead of rejecting "outlier" data that conflicts with the model, this strips out "inlier" data to get a clearer picture of how the market changes for larger moves.
Social media analytics allows us to extract, analyze, and establish semantic from user-generated contents in social media platforms. This study utilized a mixed method including a three-step process of data collection, topic modeling, and data annotation for recognizing exercise related patterns. Based on the findings,…
Biarchetype analysis identifies extreme instances of observations and features.
CEDA improves understanding of data fit to models.
As data collections become larger, exploratory regression analysis becomes more important but more challenging. When observations are hierarchically clustered the problem is even more challenging because model selection with mixed effect models can produce misleading results when nonlinear effects are not included into…
Factor analysis aims to determine latent factors, or traits, which summarize a given data set. Inter-battery factor analysis extends this notion to multiple views of the data. In this paper we show how a nonlinear, nonparametric version of these models can be recovered through the Gaussian process latent variable model…
ClusterGraph visualizes and simplifies multidimensional data clusters for better understanding.
Study predicts adverse events in Afghanistan using time series data.
Study analyzes Disney stock market performance using machine learning.
Exploratory data analysis is crucial for developing and understanding classification models from high-dimensional datasets. We explore the utility of a new unsupervised tree ensemble called uncharted forest for visualizing class associations, sample-sample associations, class heterogeneity, and uninformative classes fo…
In conventional supervised learning, a training dataset is given with ground-truth labels from a known label set, and the learned model will classify unseen instances to known labels. This paper studies a new problem setting in which there are unknown classes in the training data misperceived as other labels, and thus …
In this paper, we propose a new algorithm for exploratory projection pursuit. The basis of the algorithm is the insight that previous approaches used fairly narrow definitions of interestingness / non interestingness. We argue that allowing these definitions to depend on the problem / data at hand is a more natural app…
Marginal maximum likelihood (MML) estimation is the preferred approach to fitting item response theory models in psychometrics due to the MML estimator's consistency, normality, and efficiency as the sample size tends to infinity. However, state-of-the-art MML estimation procedures such as the Metropolis-Hastings Robbi…
Modern data is messy and high-dimensional, and it is often not clear a priori what are the right questions to ask. Instead, the analyst typically needs to use the data to search for interesting analyses to perform and hypotheses to test. This is an adaptive process, where the choice of analysis to be performed next dep…
Wasserstein t-SNE embeds hierarchical datasets considering within-unit distributions.
Python package for functional data analysis.
The paper analyzes Lending Club's loan applicants to predict default risk.
Latent feature modeling allows capturing the latent structure responsible for generating the observed properties of a set of objects. It is often used to make predictions either for new values of interest or missing information in the original data, as well as to perform data exploratory analysis. However, although the…
Proposes a flexible feature allocation model for sparse factor analysis.
Performing diagnosis or exploratory analysis during the training of deep learning models is challenging but often necessary for making a sequence of decisions guided by the incremental observations. Currently available systems for this purpose are limited to monitoring only the logged data that must be specified before…
Visual exploration of high-dimensional real-valued datasets is a fundamental task in exploratory data analysis (EDA). Existing methods use predefined criteria to choose the representation of data. There is a lack of methods that (i) elicit from the user what she has learned from the data and (ii) show patterns that she…
cGAP visualizes high-dimensional categorical data with interpretable geometric structure.
cGAP visualizes high-dimensional categorical data with interpretable geometric structure.
Yarbus' claim to decode the observer's task from eye movements has received mixed reactions. In this paper, we have supported the hypothesis that it is possible to decode the task. We conducted an exploratory analysis on the dataset by projecting features and data points into a scatter plot to visualize the nuance prop…
Study speculative trading using RL with exploratory framework.
Sampling strategies significantly affect feature approximations in ELA, impacting classifier accuracy.
New Shapley values reveal non-linear feature dependencies.
This paper analyzes the quantitative relations between stock prices and quantities of tradable stock shares in Chinese stock markets at six time points by means of Exploratory Data Analysis (EDA) method. It is found the resulting formulae have the same structure but different parameters. This paper also uses these rela…
Advances neural tri-factorization for clustering and discordance analysis of multi-typed data.
A-DOGE embeds attributed graphs efficiently using density of states.
Breaks the hardness conjecture for batch RL with a novel tournament-based approach.
Study of entropy-regularized LQG MFGs with exploratory actions.
New method visualizes noisy data better than existing techniques.
The paper tackles confidence calibration for exploratory machine learning problems.