Examines insurance market development and similarity post-2004 EU enlargement.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We develop a local theory for the construction of singular spacetimes in all spacetime dimensions which become asymptotically self-similar as the singularity is approached. The techniques developed also allow us to construct and classify exact self-similar solutions which correspond to the formal asymptotic expansions …
Turtle Score analyzes developer similarity to match high-performing candidates.
Develops an ordinal-similarity framework for scalable and interpretable representation alignment.
Paper develops active learning for clustering unknown pairwise similarities.
In this study, we define a family of ruled surfaces in the Euclidean 3-space E^3 and called similar ruled surfaces. We obtain some properties of these special surfaces and we show that developable ruled surfaces form a family of similar ruled surfaces if and only if the striction curves of the surfaces are similar curv…
In this paper we do the first large scale analysis of writing style development among Danish high school students. More than 10K students with more than 100K essays are analyzed. Writing style itself is often studied in the natural language processing community, but usually with the goal of verifying authorship, assess…
In this study, we consider the notion of similar ruled surface for timelike and spacelike ruled surfaces in Minkowski 3-space. We obtain some properties of these special surfaces in E_1^3 and we show that developable ruled surfaces in E_1^3 form a family of similar ruled surfaces if and only if the striction curves of …
Paper develops multivariate time series similarity and distance measures.
Language-based methods improve human similarity approximations without requiring many human judgments.
The problem of hierarchical clustering items from pairwise similarities is found across various scientific disciplines, from biology to networking. Often, applications of clustering techniques are limited by the cost of obtaining similarities between pairs of items. While prior work has been developed to reconstruct cl…
Paper develops a new similarity metric for predicting stock market returns.
Develops a fast method to analyze image similarities without pre-trained models.
Improves confidence calibration in neural networks by smoothing labels based on class similarity.
A Sim(n-1,1) affine manifold is an affine manifold whose linear holonomy is contained in the similarity lorentzian group but not in the lorentzian group. The class of similarity lorentzian affine manifolds is a small part in the nice class of conformally lorentzian flat manifolds. In this paper we show that a compact S…
In [9] Kaimanovich introduced the concept of augmented tree on the symbolic space of a self-similar set. It is hyperbolic in the sense of Gromov, and it was shown in [13] that under the open set condition, a self-similar set can be identified with the hyperbolic boundary of the tree. In the paper, we investigate in det…
The paper proposes using non-isotropic distances for more accurate trace link recommendation.
Modified cosine distance improves similarity performance in data with variance and correlation.
If an -manifold is locally modeled on $\RR^{m+2}$ with coordinate changes lying in the subgroup $G=\RR^{m+2}\rtimes ({\rO}(m+1,1)\times \RR^+)$ of the affine group ${\rA}(m+2)$, then is said to be a \emph{Lorentzian similarity manifold}. A Lorentzian similarity manifold is also a conformally flat Lorentzia…
CLS measures dataset similarity through decision rule performance.
BiLRP explains deep similarity models by decomposing scores into feature contributions.
The paper develops a decision support system for hierarchical text classification of conference proceedings.
We introduce a new discrepancy score between two distributions that gives an indication on their similarity. While much research has been done to determine if two samples come from exactly the same distribution, much less research considered the problem of determining if two finite samples come from similar distributio…
We develop a theory of Nobeling manifolds similar to the theory of Hilbert space manifolds. We show that it reflects the theory of Menger manifolds developed by M. Bestvina and is its counterpart in the realm of complete spaces. In particular, the Nobeling manifold characterization conjecture is proven.
We study the dynamics of correlation and variance in systems under the load of environmental factors. A universal effect in ensembles of similar systems under the load of similar factors is described: in crisis, typically, even before obvious symptoms of crisis appear, correlation increases, and, at the same time, vari…
Matrix factorization is a key component of collaborative filtering-based recommendation systems because it allows us to complete sparse user-by-item ratings matrices under a low-rank assumption that encodes the belief that similar users give similar ratings and that similar items garner similar ratings. This paradigm h…
Similarity algebra extends algebraic structures with quantitative bounds.
ROTS improves sentence similarity by incorporating structural information.
We propose a novel method of introducing structure into existing machine learning techniques by developing structure-based similarity and distance measures. To learn structural information, low-dimensional structure of the data is captured by solving a non-linear, low-rank representation problem. We show that this low-…
Despite the advances of deep learning in specific tasks using images, the principled assessment of image fidelity and similarity is still a critical ability to develop. As it has been shown that Mean Squared Error (MSE) is insufficient for this task, other measures have been developed with one of the most effective bei…
Recently, metric learning and similarity learning have attracted a large amount of interest. Many models and optimisation algorithms have been proposed. However, there is relatively little work on the generalization analysis of such methods. In this paper, we derive novel generalization bounds of metric and similarity …
The development of algorithms for hierarchical clustering has been hampered by a shortage of precise objective functions. To help address this situation, we introduce a simple cost function on hierarchies over a set of points, given pairwise similarities between those points. We show that this criterion behaves sensibl…
There are plenty of problems where the data available is scarce and expensive. We propose a generator of semi-artificial data with similar properties to the original data which enables development and testing of different data mining algorithms and optimization of their parameters. The generated data allow a large scal…
We review some developments on clustering stochastic processes and come with the conclusion that asymptotically consistent clustering algorithms can be obtained when the processes are ergodic and the dissimilarity measure satisfies the triangle inequality. Examples are provided when the processes are distribution ergod…
Twinning splits data into fast, statistically similar sets.
In a context of document co-clustering, we define a new similarity measure which iteratively computes similarity while combining fuzzy sets in a three-partite graph. The fuzzy triadic similarity (FT-Sim) model can deal with uncertainty offers by the fuzzy sets. Moreover, with the development of the Web and the high ava…
New insights into continual learning with task similarity.
Improves document summarization by combining word embeddings and n-grams.
Data similarity is a key concept in many data-driven applications. Many algorithms are sensitive to similarity measures. To tackle this fundamental problem, automatically learning of similarity information from data via self-expression has been developed and successfully applied in various models, such as low-rank repr…
A novel approach to analyzing time series generated by complex systems, such as markets, is presented. The basic idea of the approach is the {\it Law of Self-Similar Evolution}, according to which any complex system develops self-similarly. There always exist some internal laws governing the evolution of a system, say …
In this paper, we build up a min-max theory for minimal surfaces using sweepouts of surfaces of genus . We develop a direct variational methods similar to the proof of the famous Plateau problem by J. Douglas and T. Rado. As a result, we show that the min-max value for the area functional can be achieved by a …
This paper tackles efficient optimization for nonlinear embeddings in similarity learning.
The special linear groups, the mapping class groups of surfaces, the outer autormorphism groups of free groups appear in numerous domains. Their analogies, developped in particular in K. Vogtmann's work, have been written about a lot. In this report, we concentrate on the contractible spaces on which these groups act i…
New interpretation reconciles country and product complexity.
Unsupervised text embedding has shown great power in a wide range of NLP tasks. While text embeddings are typically learned in the Euclidean space, directional similarity is often more effective in tasks such as word similarity and document clustering, which creates a gap between the training stage and usage stage of t…
Background: The problem of predicting whether a drug combination of arbitrary orders is likely to induce adverse drug reactions is considered in this manuscript. Methods: Novel kernels over drug combinations of arbitrary orders are developed within support vector machines for the prediction. Graph matching methods are …
Real networks exhibit nontrivial topological features such as heavy-tailed degree distribution, high clustering, and small-worldness. Researchers have developed several generative models for synthesizing artificial networks that are structurally similar to real networks. An important research problem is to identify the…
Traditional semantic similarity models often fail to encapsulate the external context in which texts are situated. However, textual datasets generated on mobile platforms can help us build a truer representation of semantic similarity by introducing multimodal data. This is especially important in sparse datasets, maki…