Combines topological and geometric approaches to data analysis.
problem Understanding when and how geometric objects intersect.
method Connects topological and geometric concepts of curvature.
result Reconceptualizes curvature and links it to hyperconvexity.
This paper investigates to identify the requirement and the development of machine learning-based mobile big data analysis through discussing the insights of challenges in the mobile big data (MBD). Furthermore, it reviews the state-of-the-art applications of data analysis in the area of MBD. Firstly, we introduce the …
New method uses topological data analysis to study stock market crashes.
problem Characterizing and predicting stock market crashes.
method Topological data analysis, persistence landscape, dynamic time series analysis.
result Demonstrates effectiveness of new method for Flash Crash characterization and prediction.
This study analyzes data science vocabulary changes over 13 years.
problem Understanding evolution of data science terms over time.
method Exploratory Data Analysis, Latent Semantic Analysis, Latent Dirichlet Analysis, N-grams Analysis.
result Identified new vocabulary and its incorporation into scientific literature.
Study uses big data to analyze quantum invariants.
problem Investigate structural properties of Jones polynomial.
method Exploratory and topological data analysis, including coloring, rank increase, categorification.
result Contrasts behavior of Jones polynomial under various enhancements.
CEDA improves understanding of data fit to models.
problem Real-world data often deviates from theoretical models.
method Categorical Exploratory Data Analysis (CEDA) to highlight deviations.
result CEDA reveals where and how data fits or deviates from models.
In this paper the exact linear relation between the leading eigenvectors of the modularity matrix and the singular vectors of an uncentered data matrix is developed. Based on this analysis the concept of a modularity component is defined, and its properties are developed. It is shown that modularity component analysis …
The increasing availability of large but noisy data sets with a large number of heterogeneous variables leads to the increasing interest in the automation of common tasks for data analysis. The most time-consuming part of this process is the Exploratory Data Analysis, crucial for better domain understanding, data clean…
Topological data analysis quantifies structural dynamics using persistent homology.
problem Analyzing the shape and topology of structural dynamics data.
method Topological Data Analysis (TDA) with persistent homology to quantify shape over scales.
result Persistent homology reveals significant changes in manifold shape due to damage, not temperature.
Following a review of metric, ultrametric and generalized ultrametric, we review their application in data analysis. We show how they allow us to explore both geometry and topology of information, starting with measured data. Some themes are then developed based on the use of metric, ultrametric and generalized ultrame…
Proposes MVGPR for spatiotemporal data modal analysis.
problem Sparse and irregularly sampled data in complex flows.
method Multivariate Gaussian process regression (MVGPR) with kernel design.
result MVGPR outperforms DMD and SPOD in modal analysis of sparse and irregular data.
New tool helps analyze complex financial data.
problem Difficulty in comprehending high-dimensional financial data.
method Topological Data Analysis Ball Mapper algorithm.
result Shows new way to see detail in financial data.
Paper introduces topological eigenvalue theorems for tensor analysis in multi-modal data.
problem Lack of deep understanding of tensor structures in multi-modal data fusion.
method Introduces topological perspective to tensor eigenvalue analysis, linking eigenvalues to topological features.
result Establishes new theorems that enhance understanding of tensor structures in data fusion.
We present new findings in regard to data analysis in very high dimensional spaces. We use dimensionalities up to around one million. A particular benefit of Correspondence Analysis is its suitability for carrying out an orthonormal mapping, or scaling, of power law distributed data. Power law distributed data are foun…
Data science enhances knot theory by analyzing invariant relations.
problem Understanding the complex relations between knot invariants.
method Topological data analysis applied to knot theory.
result New insights into long-standing conjectures about knots.
This paper introduces compositional data analysis for financial ratios, improving industry-level analysis.
problem Statistical issues with standard financial ratios at industry level.
method Compositional data analysis techniques for financial ratios.
result Improved analysis of financial ratios using compositional data methods.
With developing of computation tools in the last years, data analysis methods to find insightful information are becoming more common among industries and researchers. This paper is the first part of the times series analysis of New England electricity price and demand to find anomaly in the data. In this paper time-se…
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
New RL algorithms correct bias in dynamic data analysis.
problem Dynamic data generation and analysis create endogeneity issues.
method Instrument variable (IV)-based reinforcement learning (RL) algorithms.
result Established theoretical properties of IV-RL algorithms.
The problem of complex data analysis is a central topic of modern statistical science and learning systems and is becoming of broader interest with the increasing prevalence of high-dimensional data. The challenge is to develop statistical models and autonomous algorithms that are able to acquire knowledge from raw dat…
FinSphere improves stock analysis quality with AI and expert-curated data.
problem Lack of objective evaluation metrics and depth in stock analysis by FinLLMs.
method Developed AnalyScore, curated Stocksis dataset, and FinSphere AI agent.
result FinSphere outperforms general and domain-specific LLMs in generating high-quality stock analysis reports.
Taxicab correspondence analysis visualizes sparse text data sets.
problem Visualization of extremely sparse contingency tables.
method Robust variant of correspondence analysis for sparse data.
result Visualized an 8265-dimensional textual data set.
For multiple multivariate data sets, we derive conditions under which Generalized Canonical Correlation Analysis (GCCA) improves classification performance of the projected datasets, compared to standard Canonical Correlation Analysis (CCA) using only two data sets. We illustrate our theoretical results with simulation…
Proposes CC-NMDF for analyzing manifold-valued data.
problem Nonlinear structure in manifold-valued data requires new analysis methods.
method Curvature-corrected nonnegative manifold data factorization (CC-NMDF) with an iterative algorithm.
result Demonstrates CC-NMDF on real-world diffusion tensor MRI data.
New method simplifies data analysis.
problem Complex data analysis challenges.
method Innovative algorithm for data simplification.
result Significant reduction in analysis time.
Paper introduces RKHM and KME for richer data analysis.
problem Lack of rich data structures in kernel methods.
method Proposes RKHM and KME for functional data analysis.
result RKHM captures structural properties in functional data.
PCA simplifies multivariate extreme data analysis.
problem Analyzing multivariate extreme values with high-dimensional data.
method Principal Component Analysis (PCA) for dimensionality reduction.
result PCA helps preserve essential information for extreme value analysis.
The paper uses data science to predict stock trends of Amazon, Apple, Google, and Microsoft.
problem Short-term market movement prediction for major tech stocks.
method Combination of technical analysis and machine/deep learning for trend classification.
result Generated labels for data set: +1 (buy), 0 (hold), -1 (sell).
Novel method converts time series data into functional data for high dimensional classification.
problem Small sample size problem in high dimensional time series data.
method Classwise Functional Principal Component Analysis (PCA) followed by Bayesian linear classifier.
result Demonstrated efficacy on synthetic and real data sets.
R-PCA extends PCA to Riemannian manifolds for structured data.
problem Applying PCA to data on Riemannian manifolds without vector space operations.
method Adapting PCA to Riemannian manifolds by equipping data with local metrics.
result Unified approach for dimensionality reduction and statistical analysis on manifolds.
Research in several fields now requires the analysis of data sets in which multiple high-dimensional types of data are available for a common set of objects. In particular, The Cancer Genome Atlas (TCGA) includes data from several diverse genomic technologies on the same cancerous tumor samples. In this paper we introd…
This paper quantifies privacy loss in exploratory data analysis.
problem Privacy loss in exploratory data analysis is often overlooked in privacy budgets.
method Quantitative analysis of privacy loss for statistical functions.
result Privacy loss must be considered in calculating machine learning privacy budgets.
A new robust scaling approach improves downstream metabolomics analysis.
problem Challenges in choosing scaling techniques for metabolomics data.
method Introduces a weighted scaling approach robust to outliers.
result The proposed method outperforms traditional scaling techniques in both outlier-free and outlier-present datasets.
Prototypal analysis is introduced to overcome two shortcomings of archetypal analysis: its sensitivity to outliers and its non-locality, which reduces its applicability as a learning tool. Same as archetypal analysis, prototypal analysis finds prototypes through convex combination of the data points and approximates th…
Topological data analysis classifies encrypted bits with success.
problem Classifying encrypted data with traditional machine learning methods.
method Persistent homology for generating topological features, machine learning pipeline.
result Successfully classifies encrypted data, outperforming classical models.
This work develops methods to analyze data on curved spaces using deep learning.
problem Analyzing data in non-linear, curved spaces.
method Pullback Riemannian geometry through diffeomorphisms.
result Diffeomorphisms need to map data into geodesic subspaces to ensure proper data analysis.
Can textual data be compressed intelligently without losing accuracy in evaluating sentiment? In this study, we propose a novel evolutionary compression algorithm, PARSEC (PARts-of-Speech for sEntiment Compression), which makes use of Parts-of-Speech tags to compress text in a way that sacrifices minimal classification…
New method analyzes knots and links using multiscale Gauss link integral.
problem Lack of localization and quantization in knot theory applications.
method Integrates curve segmentation and multiscale analysis into the Gauss link integral.
result Significantly outperforms other methods in protein flexibility analysis.
Paper combines geometry and time-series analysis for spatiotemporal data.
problem Multivariate time-series data from multiple sensors.
method Combines manifold learning, Riemannian geometry, and spectral analysis.
result Proposes Riemannian multi-resolution analysis (RMRA) for dynamic mode extraction.
Survey of de Casteljau's algorithm's applications in geometric data analysis.
problem No specific problem stated; focuses on algorithm applications.
method Constructive approach to generalize parametric smooth curves to manifolds.
result Algorithm provides principled way to analyze geometric data.
Paper introduces probabilistic methods to approximate archetypal analysis, reducing complexity.
problem Inherent computational complexity of archetypal analysis limits its practical applicability.
method Two preprocessing techniques: dimensionality reduction and representation cardinality reduction, using probabilistic geometry.
result The method effectively reduces scaling and provides near-optimal solutions for prediction errors.
New robust discriminant analysis for non-Gaussian data.
problem Classical discriminant analysis struggles with non-Gaussian distributions and contaminated datasets.
method Each data point follows its own ES distribution with arbitrary scale, leading to robust classification.
result Maximum-likelihood estimation and classification are simple, fast, and robust.
A new method for LDA using randomized Kaczmarz improves accuracy for large datasets.
problem Efficiently performing LDA on large datasets.
method Randomized Kaczmarz method applied to linear discriminant analysis.
result The method achieves comparable accuracy to full data LDA.
The paper gives picture of enrichment to economic and financial system analysis using agent-based models as a form of advanced study for financial economic data post-statistical-data analysis and micro-simulation analysis. Theoretical exploration is carried out by using comparisons of some usual financial economy syste…
Intelligent financial data analysis system improves accuracy and efficiency.
problem Inefficient and inaccurate financial data analysis due to complex data and evolving contexts.
method Integrates LLMs with RAG technology for financial data analysis.
result Significant improvements in accuracy and recall (78.6% and 89.2%) compared to baseline.
New method for analyzing multiple longitudinal data processes.
problem Exploring associations between multiple random processes observed jointly.
method Functional Generalized Canonical Correlation Analysis (FGCCA) based on multiblock Regularized Generalized Canonical Correlation Analysis (RGCCA).
result FGCCA framework is robust to sparsely and irregularly observed data.
Probabilistic techniques are central to data analysis, but different approaches can be difficult to apply, combine, and compare. This paper introduces composable generative population models (CGPMs), a computational abstraction that extends directed graphical models and can be used to describe and compose a broad class…
Independent Component Analysis (ICA) - one of the basic tools in data analysis - aims to find a coordinate system in which the components of the data are independent. In this paper we present Multiple-weighted Independent Component Analysis (MWeICA) algorithm, a new ICA method which is based on approximate diagonalizat…