BERTino is a lightweight Italian DistilBERT model for NLP tasks.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Word2Vec embedding for Italian language developed.
This dataset contains the annual aggregated income taxes of all the Italian municipalities over the years 2007-2011. Data are clustered over the Italian regions and provinces. The source of the data is the Italian Ministry of Economics and Finance. The administrative variations in Italy over the quinquennium have been …
Many methods have been used to recognize author personality traits from text, typically combining linguistic feature engineering with shallow learning models, e.g. linear regression or Support Vector Machines. This work uses deep-learning-based models and atomic features of text, the characters, to build hierarchical, …
A framework for flagging content with limited data.
The global financial crisis, beginning in 2008, took an historic toll on national economies around the world. Following equity market crashes, unemployment rates rose significantly in many countries: Italy was among those. What will be the impact of such large shocks on Italian healthcare finances? An empirical model f…
Recently, sentiment analysis has received a lot of attention due to the interest in mining opinions of social media users. Sentiment analysis consists in determining the polarity of a given text, i.e., its degree of positiveness or negativeness. Traditionally, Sentiment Analysis algorithms have been tailored to a speci…
One of the main issues affecting the Italian NHS is the healthcare deficit: according to current agreements between the Italian State and its Regions, public funding of regional NHS is now limited to the amount of regional deficit and is subject to previous assessment of strict adherence to constraint on regional healt…
This paper improves speech recognition by distilling knowledge from acoustic models.
Study predicts firm defaults using machine learning on Italian credit data.
We present NN-grams, a novel, hybrid language model integrating n-grams and neural networks (NN) for speech recognition. The model takes as input both word histories as well as n-gram counts. Thus, it combines the memorization capacity and scalability of an n-gram model with the generalization ability of neural network…
In this paper, we empirically study models for pricing Italian sovereign bonds under a reduced form framework, by assuming different dynamics for the short-rate process. We analyze classical Cox-Ingersoll-Ross and Vasicek multi-factor models, with a focus on optimization algorithms applied in the calibration exercise. …
Covid lockdown increased interest in Italian stock market, leading to new investors.
Using a data set which includes all transactions among banks in the Italian money market, we study their trading strategies and the dependence among them. We use the Fourier method to compute the variance-covariance matrix of trading strategies. Our results indicate that well defined patterns arise. Two main communitie…
We investigate the shape of the Italian personal income distribution using microdata from the Survey on Household Income and Wealth, made publicly available by the Bank of Italy for the years 1977--2002. We find that the upper tail of the distribution is consistent with a Pareto-power law type distribution, while the r…
Italian banks use swaps to hedge against rising interest rates, offsetting losses on debt securities.
Italy and the Eurozone are heading in the year 2012 into a financial depression of unprecedented magnitude, with a forthcoming multitude of often contradictory public economic and financial stability emergency interventions whose ultimate endogenous and exogenous effects on public and private health spending and on the…
This is my master thesis. Unfortunately it is written in Italian, but maybe somebody will find it helpful when it comes to Evans Potentials
AI uses KGs to assess economic impact of selective lockdowns on Italian companies.
This paper explores a real-world fundamental theme under a data science perspective. It specifically discusses whether fraud or manipulation can be observed in and from municipality income tax size distributions, through their aggregation from citizen fiscal reports. The study case pertains to official data obtained fr…
We study the gap between the state pension provided by the Italian pension system pre-Dini reform and post-Dini reform. The goal is to fill the gap between the old and the new pension by joining a defined contribution pension scheme and adopting an optimal investment strategy that is target-based. We find that it is po…
The yearly aggregated tax income data of all, more than 8000, Italian municipalities are analyzed for a period of five years, from 2007 to 2011, to search for conformity or not with Benford's law, a counter-intuitive phenomenon observed in large tabulated data where the occurrence of numbers having smaller initial digi…
In recent years, the interest in Big Data sources has been steadily growing within the Official Statistic community. The Italian National Institute of Statistics (Istat) is currently carrying out several Big Data pilot studies. One of these studies, the ICT Big Data pilot, aims at exploiting massive amounts of textual …
In this paper we describe three stochastic models based on a semi-Markov chains approach and its generalizations to study the high frequency price dynamics of traded stocks. The three models are: a simple semi-Markov chain model, an indexed semi-Markov chain model and a weighted indexed semi-Markov chain model. We show…
The article describes the algorithm used to define the electricity price in day-ahead and itraday energy markets in Italy. Details of Matlab implementation of one of its simplified versions, capable of producing good results in a extremely short time, are then provided and numerical results are discussed.
The paper predicts financial markets using news text and semantic network analysis.
The present work constitutes the second part of a two-paper project that, in particular, deals with an in-depth study of effective techniques used in econometrics in order to make accurate forecasts in the concrete framework of one of the major economies of the most productive Italian area, namely the province of Veron…
This study uses Tsallis entropy to analyze diversification and integration in Italian stock market companies.
Background. In Italy, in recent years, vaccination coverage for key immunizations as MMR has been declining to worryingly low levels. In 2017, the Italian Gov't expanded the number of mandatory immunizations introducing penalties to unvaccinated children's families. During the 2018 general elections campaign, immunizat…
Paper uses ML to predict SME defaults with interpretability.
We analyze the data of the Italian and U.S. futures on the stock markets and we test the validity of the Continuous Time Random Walk assumption for the survival probability of the returns time series via a renewal aging experiment. We also study the survival probability of returns sign and apply a coarse graining proce…
Tax evasion is the illegal evasion of taxes by individuals, corporations, and trusts. The revenue loss from tax avoidance can undermine the effectiveness and equity of the government policies. A standard measure of tax evasion is the tax gap, that can be estimated as the difference between the total amounts of tax theo…
This paper presents a simple model to measure the relative economic growth of economic systems. The model considers S-Shaped patterns of economic growth that, represented with a linear model, measure how an economic system grows in comparison with another one. In particular, this model introduces an approach which indi…
Gas demand forecasting is a critical task for energy providers as it impacts on pipe reservation and stock planning. In this paper, the one-day-ahead forecasting of residential gas demand at country level is investigated by implementing and comparing five models: Ridge Regression, Gaussian Process (GP), k-Nearest Neigh…
This research presents an analysis of the demographic risk related to future membership patterns in pension funds with restricted entrance, financed under a pay-as-you-go scheme. The paper, therefore, proposes a stochastic model for investigating the behaviour of the demographic variable "new entrants" and the influenc…
This paper focuses on a comparative evaluation of the most common and modern methods for text classification, including the recent deep learning strategies and ensemble methods. The study is motivated by a challenging real data problem, characterized by high-dimensional and extremely sparse data, deriving from incoming…
Instabilities in the price dynamics of a large number of financial assets are a clear sign of systemic events. By investigating a set of 20 high cap stocks traded at the Italian Stock Exchange, we find that there is a large number of high frequency cojumps. We show that the dynamics of these jumps is described neither …
In this paper we propose a new stochastic model based on a generalization of semi-Markov chains to study the high frequency price dynamics of traded stocks. We assume that the financial returns are described by a weighted indexed semi-Markov chain model. We show, through Monte Carlo simulations, that the model is able …
The Wallenius distribution is a generalisation of the Hypergeometric distribution where weights are assigned to balls of different colours. This naturally defines a model for ranking categories which can be used for classification purposes. Since, in general, the resulting likelihood is not analytically available, we a…
In this paper we propose a bivariate generalization of a weighted indexed semi-Markov chains to study the high frequency price dynamics of traded stocks. We assume that financial returns are described by a weighted indexed semi-Markov chain model. We show, through Monte Carlo simulations, that the model is able to repr…
In this paper we study the high frequency dynamic of financial volumes of traded stocks by using a semi-Markov approach. More precisely we assume that the intraday logarithmic change of volume is described by a weighted-indexed semi-Markov chain model. Based on this assumptions we show that this model is able to reprod…
We study the high frequency price dynamics of traded stocks by a model of returns using a semi-Markov approach. More precisely we assume that the intraday returns are described by a discrete time homogeneous semi-Markov which depends also on a memory index. The index is introduced to take into account periods of high a…
Relationship lending is broadly interpreted as a strong partnership between a lender and a borrower. Nevertheless, we still lack consensus regarding how to quantify the strength of a lending relationship, while simple statistics such as the frequency and volume of loans have been used as proxies in previous studies. He…
We study the high frequency price dynamics of traded stocks by a model of returns using a semi-Markov approach. More precisely we assume that the intraday return are described by a discrete time homogeneous semi-Markov process and the overnight returns are modeled by a Markov chain. Based on this assumptions we derived…
In this work, we consider Corporate Governance (CG) ties among companies from a multiple network perspective. Such a structure naturally arises from the close interrelation between the Shareholding Network (SH) and the Board of Directors network (BD). In order to capture the simultaneous effects of both networks on CG,…
We study the volatility of the MIB30-stock-index high-frequency data from November 28, 1994 through September 15, 1995. Our aim is to empirically characterize the volatility random walk in the framework of continuous-time finance. To this end, we compute the index volatility by means of the log-return standard deviatio…
The number of Italian firms in function of the number of workers is well approximated by an inverse power law up to 15 workers but shows a clear downward deflection beyond this point, both when using old pre-1999 data and when using recent (2014) data. This phenomenon could be associated with employent protection legis…
We use the theory of complex networks in order to quantitatively characterize the formation of communities in a particular financial market. The system is composed by different banks exchanging on a daily basis loans and debts of liquidity. Through topological analysis and by means of a model of network growth we can d…