Bibliographic analysis considers author's research areas, the citation network and paper content among other things. In this paper, we combine these three in a topic model that produces a bibliographic model of authors, topics and documents using a non-parametric extension of a combination of the Poisson mixed-topic li…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Bibliographic analysis considers the author's research areas, the citation network and the paper content among other things. In this paper, we combine these three in a topic model that produces a bibliographic model of authors, topics and documents, using a nonparametric extension of a combination of the Poisson mixed-…
FONDUE identifies ambiguous nodes in networks for better analysis.
I describe the history of Topological Tverberg Theorem. I present some important constructions and discuss their properties. In particular, I describe in details the cell structure of the classifying space where is the permutation group. I also clarify some bibliographical issues.
Felix Klein's so-called Erlangen Program was published in 1872 as professoral dissertation. It proposed a new solution to the problem how to classify and characterize geometries on the basis of projective geometry and group theory. The given translation was made in 1892 by Dr. M. W. Haskell and transcribed by N. C. Rug…
In this work, we consider the problem of combining link, content and temporal analysis for community detection and prediction in evolving networks. Such temporal and content-rich networks occur in many real-life settings, such as bibliographic networks and question answering forums. Most of the work in the literature (…
Blind source separation (BSS), i.e., the decoupling of unknown signals that have been mixed in an unknown way, has been a topic of great interest in the signal processing community for the last decade, covering a wide range of applications in such diverse fields as digital communications, pattern recognition, biomedica…
Systematic review of ML explainability in process mining.
Author name disambiguation in bibliographic databases is the problem of grouping together scientific publications written by the same person, accounting for potential homonyms and/or synonyms. Among solutions to this problem, digital libraries are increasingly offering tools for authors to manually curate their publica…
This article reviews entity resolution methods and their applications.
In supervised machine learning for author name disambiguation, negative training data are often dominantly larger than positive training data. This paper examines how the ratios of negative to positive training data can affect the performance of machine learning algorithms to disambiguate author names in bibliographic …
Study models live cattle futures prices in Brazil.
Background: Overweight and obesity are an increasing phenomenon worldwide. Predicting future overweight or obesity early in the childhood reliably could enable a successful intervention by experts. While a lot of research has been done using explanatory modeling methods, capability of machine learning, and predictive m…
This book, which is in Spanish, provides detailed descriptions, including over 550 mathematical formulas, for over 150 trading strategies across a host of asset classes (and trading styles). This includes stocks, options, fixed income, futures, ETFs, indexes, commodities, foreign exchange, convertibles, structured asse…
Efficient private matrix analysis algorithms for recent variants.
This paper introduces compositional data analysis for financial ratios, improving industry-level analysis.
In this dissertation, the main goal is visualisation of financial time series. We expect that visualisation of financial time series will be a useful auxiliary for technical analysis. Firstly, we review the technical analysis methods and test our trading rules, which are built by the essential concepts of technical ana…
Paper combines geometry and time-series analysis for spatiotemporal data.
In this paper the exact linear relation between the leading eigenvectors of the modularity matrix and the singular vectors of an uncentered data matrix is developed. Based on this analysis the concept of a modularity component is defined, and its properties are developed. It is shown that modularity component analysis …
This paper investigates to identify the requirement and the development of machine learning-based mobile big data analysis through discussing the insights of challenges in the mobile big data (MBD). Furthermore, it reviews the state-of-the-art applications of data analysis in the area of MBD. Firstly, we introduce the …
Interactive DR framework for comparing datasets.
Combines topological and geometric approaches to data analysis.
Proposes a multivariate regression model for better analysis of multiple datasets.
New method uses topological data analysis to study stock market crashes.
Analyzes stock trends and e-commerce user behavior using Twitter data.
Study analyzes Disney stock market performance using machine learning.
This study analyzes data science vocabulary changes over 13 years.
FinSphere improves stock analysis quality with AI and expert-curated data.
This study examines whether PCA can effectively identify nitrogen pollution sources in rivers.
Paper introduces probabilistic methods to approximate archetypal analysis, reducing complexity.
We present a unifying framework which reduces the construction of probabilistic component analysis techniques to a mere selection of the latent neighbourhood, thus providing an elegant and principled framework for creating novel component analysis models as well as constructing probabilistic equivalents of deterministi…
Deep learning techniques are rapidly advanced recently, and becoming a necessity component for widespread systems. However, the inference process of deep learning is black-box, and not very suitable to safety-critical systems which must exhibit high transparency. In this paper, to address this black-box limitation, we …
In this paper, we explore the effectiveness of dynamic analysis techniques for identifying malware, using Hidden Markov Models (HMMs) and Profile Hidden Markov Models (PHMMs), both trained on sequences of API calls. We contrast our results to static analysis using HMMs trained on sequences of opcodes, and show that dyn…
Improved LDA method for better classification and dimensionality reduction.
CDVI improves variational inference for survival analysis by considering censoring mechanisms.
The kernel matrix used in kernel methods encodes all the information required for solving complex nonlinear problems defined on data representations in the input space using simple, but implicitly defined, solutions. Spectral analysis on the kernel matrix defines an explicit nonlinear mapping of the input data represen…
Prototypal analysis is introduced to overcome two shortcomings of archetypal analysis: its sensitivity to outliers and its non-locality, which reduces its applicability as a learning tool. Same as archetypal analysis, prototypal analysis finds prototypes through convex combination of the data points and approximates th…
In this paper, we introduce a Deep Convolutional Analysis Dictionary Model (DeepCAM) by learning convolutional dictionaries instead of unstructured dictionaries as in the case of deep analysis dictionary model introduced in the companion paper. Convolutional dictionaries are more suitable for processing high-dimensiona…
Improved analysis of UCBVI algorithm with better empirical performance.
Reduces survival analysis to common regression tasks.
A new robust scaling approach improves downstream metabolomics analysis.
Paper proposes FinAR-Bench to evaluate LLMs in financial analysis tasks.
Wavelet analysis reveals non-linear dynamics in cryptocurrency prices.
The book explores alternatives to worst-case analysis for algorithm performance.
Alternative wavelet analysis method for financial signals.
These are lecture notes for the course "Analysis and X-ray tomography". The course is a broad overview of various tools in analysis that can be used to study X-ray tomography. The focus is on tools and ideas, not so much on technical details and minimal assumptions. Only very basic functional analysis is assumed as bac…
We investigate the difference between using an penalty versus an constraint in generalized eigenvalue problems, such as principal component analysis and discriminant analysis. Our main finding is that an penalty may fail to provide very sparse solutions; a severe disadvantage for variable sel…
Can textual data be compressed intelligently without losing accuracy in evaluating sentiment? In this study, we propose a novel evolutionary compression algorithm, PARSEC (PARts-of-Speech for sEntiment Compression), which makes use of Parts-of-Speech tags to compress text in a way that sacrifices minimal classification…