Validation is one of the most important aspects of clustering, but most approaches have been batch methods. Recently, interest has grown in providing incremental alternatives. This paper extends the incremental cluster validity index (iCVI) family to include incremental versions of Calinski-Harabasz (iCH), I index and …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Cluster analysis is used to explore structure in unlabeled data sets in a wide range of applications. An important part of cluster analysis is validating the quality of computationally obtained clusters. A large number of different internal indices have been developed for validation in the offline setting. However, thi…
iCVI-ARTMAP accelerates clustering with adaptive resonance theory and validity indices.
Tensor decompositions are invaluable tools in analyzing multimodal datasets. In many real-world scenarios, such datasets are far from being static, to the contrary they tend to grow over time. For instance, in an online social network setting, as we observe new interactions over time, our dataset gets updated in its "t…
Recent studies have found that the log-volatility of asset returns exhibit roughness. This study investigates roughness or the anti-persistence of Bitcoin volatility. Using the multifractal detrended fluctuation analysis, we obtain the generalized Hurst exponent of the log-volatility increments and find that the genera…
The statistical properties of the increments x(t+T) - x(t) of a financial time series depend on the time resolution T on which the increments are considered. A non-parametric approach is used to study the scale dependence of the empirical distribution of the price increments x(t+T) - x(t) of S&P Index futures, for time…
This paper presents a novel approach for incremental semiparametric inverse dynamics learning. In particular, we consider the mixture of two approaches: Parametric modeling based on rigid body dynamics equations and nonparametric modeling based on incremental kernel methods, with no prior information on the mechanical …
Mirror flows converge to a limiting flow with a convex potential.
Cross-validation (CV) is one of the main tools for performance estimation and parameter tuning in machine learning. The general recipe for computing CV estimate is to run a learning algorithm separately for each CV fold, a computationally expensive process. In this paper, we propose a new approach to reduce the computa…
Paper presents a method for imputing and forecasting structural response from incomplete sensor data.
Simulation workflow is a top-level model for the design and control of simulation process. It connects multiple simulation components with time and interaction restrictions to form a complete simulation system. Before the construction and evaluation of the component models, the validation of upper-layer simulation work…
PrIU optimizes machine learning model updates after data cleaning.
A new permutation method improves two-sample testing power.
Narrative disclosures in 10-K filings improve bankruptcy prediction beyond accounting ratios.
New algorithm improves on EM for streaming data, outperforming existing methods.
Growth rate of real GDP per capita is represented as a sum of two components -- a monotonically decreasing economic trend and fluctuations related to a specific age population change. The economic trend is modeled by an inverse function of real GDP per capita with a numerator potentially constant for the largest develo…
Robustly detects jumps in high-frequency CIR and CKLS models.
Incremental learning suffers from two challenging problems; forgetting of old knowledge and intransigence on learning new knowledge. Prediction by the model incrementally learned with a subset of the dataset are thus uncertain and the uncertainty accumulates through the tasks by knowledge transfer. To prevent overfitti…
Optimizes subset selection in sparse learning problems.
Paper improves privacy-preserving measurement of advertising incrementality.
This paper reviews and proposes a new approach for evaluating internal cluster validation indices.
Efficient algorithms solve large-scale DRSVM problems.
Various health-care applications such as assisted living, fall detection etc., require modeling of user behavior through Human Activity Recognition (HAR). HAR using mobile- and wearable-based deep learning algorithms have been on the rise owing to the advancements in pervasive computing. However, there are two other ch…
In this paper we present an incremental variant of the Twin Support Vector Machine (TWSVM) called Fuzzy Bounded Twin Support Vector Machine (FBTWSVM) to deal with large datasets and learning from data streams. We combine the TWSVM with a fuzzy membership function, so that each input has a different contribution to each…
Generative replay improves continual learning by using generated data as negative examples.
The paper challenges the validity of cluster validity measures in unsupervised learning.
Develops a new cluster validity index to find multiple optimal cluster numbers.
Introduces BCVI, a Bayesian cluster validity index for better cluster selection.
New indices for determining cluster compactness and separability.
Communication through e-mails remains to be highly formalized, conventional and indispensable method for the exchange of information over the Internet. An ever-increasing ratio and adversary nature of spam e-mails have posed a great many challenges such as uneven class distribution, unequal error cost, frequent change …
The growth rate of real GDP per capita in the biggest OECD countries is represented as a sum of two components - a steadily decreasing trend and fluctuations related to the change in some specific age population. The long term trend in the growth rate is modelled by an inverse function of real GDP per capita with a con…
A plain well-trained deep learning model often does not have the ability to learn new knowledge without forgetting the previously learned knowledge, which is known as catastrophic forgetting. Here we propose a novel method, SupportNet, to efficiently and effectively solve the catastrophic forgetting problem in the clas…
As we enter into the big data age and an avalanche of images have become readily available, recognition systems face the need to move from close, lab settings where the number of classes and training data are fixed, to dynamic scenarios where the number of categories to be recognized grows continuously over time, as we…
A new method for real-time anomaly detection in flight data.
We introduce incremental variational inference and apply it to latent Dirichlet allocation (LDA). Incremental variational inference is inspired by incremental EM and provides an alternative to stochastic variational inference. Incremental LDA can process massive document collections, does not require to set a learning …
Motivated by the need for accurate frequency information, a novel algorithm for estimating the fundamental frequency and its rate of change in three-phase power systems is developed. This is achieved through two stages of Kalman filtering. In the first stage a quaternion extended Kalman filter, which provides a unified…
A new framework uses an Incremental Transformer to design geopolymer mixtures efficiently.
Develops platforms to analyze social media data for human behavior and emotions.
The concepts of scale invariance, self-similarity and scaling have been fruitfully applied to the study of price fluctuations in financial markets. After a brief review of the properties of stable Levy distributions and their applications to market data we indicate the shortcomings of such models and describe the trunc…
DPGIIL clusters structural anomalies using transmissibility functions with deep learning and Dirichlet process.
A new framework for generating predictive features in noisy multivariate time series.
Inverse reinforcement learning (IRL) is the problem of learning the preferences of an agent from the observations of its behavior on a task. While this problem has been well investigated, the related problem of {\em online} IRL---where the observations are incrementally accrued, yet the demands of the application often…
Support Vector Data Description (SVDD) is a popular outlier detection technique which constructs a flexible description of the input data. SVDD computation time is high for large training datasets which limits its use in big-data process-monitoring applications. We propose a new iterative sampling-based method for SVDD…
New method improves online nonparametric estimators with minimal extra computation.
This paper tackles deep clustering evaluation challenges in high-dimensional data.
Deep learning has been widely accepted as a promising solution for medical image segmentation, given a sufficiently large representative dataset of images with corresponding annotations. With ever increasing amounts of annotated medical datasets, it is infeasible to train a learning method always with all data from scr…
Paper proposes faster incremental subclass discriminant analysis.
The extreme event statistics plays a very important role in the theory and practice of time series analysis. The reassembly of classical theoretical results is often undermined by non-stationarity and dependence between increments. Furthermore, the convergence to the limit distributions can be slow, requiring a huge am…