Thanks to the growing availability of spoofing databases and rapid advances in using them, systems for detecting voice spoofing attacks are becoming more and more capable, and error rates close to zero are being reached for the ASVspoof2015 database. However, speech synthesis and voice conversion paradigms that are not…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This work disentangles speech and non-speech components from found data.
A method for predicting survival using neural networks for both continuous and discrete time.
Labor productivity was studied at the microscopic level in terms of distributions based on individual firm financial data from Japan and the US. A power-law distribution in terms of firms and sector productivity was found in both countries' data. The labor productivities were not equal for nation and sectors, in contra…
We study the long-term memory in diverse stock market indices and foreign exchange rates using the Detrended Fluctuation Analysis(DFA). For all daily and high-frequency market data studied, no significant long-term memory property is detected in the return series, while a strong long-term memory property is found in th…
Survey of de Casteljau's algorithm's applications in geometric data analysis.
We investigate coresets - succinct, small summaries of large data sets - so that solutions found on the summary are provably competitive with solution found on the full data set. We provide an overview over the state-of-the-art in coreset construction for machine learning. In Section 2, we present both the intuition be…
We investigate relationship between annual electric power consumption per capita and gross domestic production (GDP) per capita for 131 countries. We found that the relationship can be fitted with a power-law function. We examine the relationship for 47 prefectures in Japan. Furthermore, we investigate values of annual…
Machine learning detects new minerals at Mars rover landing sites.
Firm size data usually do not show the normality that is often assumed in statistical analysis such as regression analysis. In this study we focus on two firm size data: the number of employees and sale. Those data deviate considerably from a normal distribution. To improve the normality of those data we transform them…
Data science teams collaborate extensively, using various tools and stakeholders.
Real estate appraisal is a complex and important task, that can be made more precise and faster with the help of automated valuation tools. Usually the value of some property is determined by taking into account both structural and geographical characteristics. However, while geographical information is easily found, o…
Using the tools developed for statistical physics, we simultaneously analyze statistical properties of the Jakarta and Kuala Lumpur Stock Exchange indices. In spite of the small number of data used in the analysis, the result shows the universal behavior of complex systems previously found in the leading stock indices.…
Unique global solutions found for specific initial data.
Ensemble learning that can be used to combine the predictions from multiple learners has been widely applied in pattern recognition, and has been reported to be more robust and accurate than the individual learners. This ensemble logic has recently also been more applied in feature selection. There are basically two st…
This thesis focuses on gaining linguistic insights into textual discussions on a word level. It was of special interest to distinguish messages that constructively contribute to a discussion from those that are detrimental to them. Thereby, we wanted to determine whether "I"- and "You"-messages are indicators for eithe…
The determination of acceptability prices of contingent claims requires the choice of a stochastic model for the underlying asset price dynamics. Given this model, optimal bid and ask prices can be found by stochastic optimization. However, the model for the underlying asset price process is typically based on data and…
Deep neural networks progressively transform their inputs across multiple processing layers. What are the geometrical properties of the representations learned by these networks? Here we study the intrinsic dimensionality (ID) of data-representations, i.e. the minimal number of parameters needed to describe a represent…
This work explores the ability of collective matrix factorization models in recommender systems to make predictions about users and items for which there is side information available but no feedback or interactions data, and proposes a new formulation with a faster cold-start prediction formula that can be used in rea…
We investigated distributions of short term price trends for high frequency stock market data. A number of trends as a function of their lengths was measured. We found that such a distribution does not fit to results following from an uncorrelated stochastic process. We proposed a simple model with a memory that gives …
We introduce a mathematical criterion defining the bubbles or the crashes in financial market price fluctuations by considering exponential fitting of the given data. By applying this criterion we can automatically extract the periods in which bubbles and crashes are identified. From stock market data of so-called the …
Rehabilitation training is the primary intervention to improve motor recovery after stroke, but a tool to measure functional training does not currently exist. To bridge this gap, we previously developed an approach to classify functional movement primitives using wearable sensors and a machine learning (ML) algorithm.…
ptype infers data types robustly in real-world data.
Paper reviews and compares methods for handling imbalanced data.
A connection is made between the Krammer representation and the Birman-Murakami-Wenzl algebra. Inspired by a dimension argument, a basis is found for a certain irrep of the algebra, and relations which generate the matrices are found. Following a rescaling and change of parameters, the matrices are found to be identica…
Autoencoder neural network is implemented to estimate the missing data. Genetic algorithm is implemented for network optimization and estimating the missing data. Missing data is treated as Missing At Random mechanism by implementing maximum likelihood algorithm. The network performance is determined by calculating the…
The original research question here is given by marketers in general, i.e., how to explain the changes in the desired timescale of the market. Tangled String, a sequence visualization tool based on the metaphor where contexts in a sequence are compared to tangled pills in a string, is here extended and diverted to dete…
We report on the occurrence of an anomaly in the price impacts of small transaction volumes following a change in the fee structure of an electronic market. We first review evidence for the existence of a master curve for price impact on the Johannesburg Stock Exchange (JSE). On attempting to re-estimate a master curve…
This paper computes Kakimizu complexes for all 11 crossing prime alternating knots.
A fast method for finding counterfactual explanations for decision forests.
New duality found linking neural network weights and activities for better generalization.
Data driven soft sensor design has recently gained immense popularity, due to advances in sensory devices, and a growing interest in data mining. While partial least squares (PLS) is traditionally used in the process literature for designing soft sensors, the statistical literature has focused on sparse learners, such …
New framework selects high-quality pretraining data without training LLMs.
We present new findings in regard to data analysis in very high dimensional spaces. We use dimensionalities up to around one million. A particular benefit of Correspondence Analysis is its suitability for carrying out an orthonormal mapping, or scaling, of power law distributed data. Power law distributed data are foun…
We present an analysis of oil prices in US$ and in other major currencies that diagnoses unsustainable faster-than-exponential behavior. This supports the hypothesis that the recent oil price run-up has been amplified by speculative behavior of the type found during a bubble-like expansion. We also attempt to unravel t…
The yearly aggregated tax income data of all, more than 8000, Italian municipalities are analyzed for a period of five years, from 2007 to 2011, to search for conformity or not with Benford's law, a counter-intuitive phenomenon observed in large tabulated data where the occurrence of numbers having smaller initial digi…
The paper studies estimation of parameters of diffusion market models from historical data. The standard definition of implied volatility for these models presents its value as an implicit function of several parameters, including the risk-free interest rate. In reality, the risk free interest rate is unknown and need …
AMPL is a new software pipeline for drug discovery models.
We obtain all possible solutions of a 1/4 Bogomol'nyi-Prasad-Sommerfield equation exactly, containing configurations made of walls, vortices and monopoles in the Higgs phase. We use supersymmetric U(N_C) gauge theories with eight supercharges with N_F fundamental hypermultiplets in the strong coupling limit. The moduli…
Bitcoin volatility can be predicted from price and alternative data.
We analyze the European transition economies and show that time series for most of major indices exhibit (i) power-law correlations in their values, power-law correlations in their magnitudes, and (iii) asymmetric probability distribution. We propose a stochastic model that can generate time series with all the previou…
Geography effect is investigated for the Chinese stock market including the Shanghai and Shenzhen stock markets, based on the daily data of individual stocks. The Shanghai city and the Guangdong province can be identified in the stock geographical sector. By investigating a geographical correlation on a geographical pa…
A method for clustering small datasets in high dimensions using random projections.
Statistical estimates can often be improved by fusion of data from several different sources. One example is so-called ensemble methods which have been successfully applied in areas such as machine learning for classification and clustering. In this paper, we present an ensemble method to improve community detection by…
New types of Delaunay hypersurfaces found in spheres.
Bipartite data is common in data engineering and brings unique challenges, particularly when it comes to clustering tasks that impose on strong structural assumptions. This work presents an unsupervised method for assessing similarity in bipartite data. Similar to some co-clustering methods, the method is based on regu…
Study examines how COVID-19 vaccine companies' popularity affects their stock prices.
Point patterns are sets or multi-sets of unordered elements that can be found in numerous data sources. However, in data analysis tasks such as classification and novelty detection, appropriate statistical models for point pattern data have not received much attention. This paper proposes the modelling of point pattern…