Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

6.3%12.5%18.8%25.0% · Apr 199319922001200920172026
48 results for developer similarity

Examines insurance market development and similarity post-2004 EU enlargement.

problem Comparing insurance markets of EU old and new members post-enlargement.
method Analyzes data from 2004 to present to compare insurance markets.
result Identifies similarities and differences in insurance markets post-2004 enlargement.

In this paper we do the first large scale analysis of writing style development among Danish high school students. More than 10K students with more than 100K essays are analyzed. Writing style itself is often studied in the natural language processing community, but usually with the goal of verifying authorship, assess…

2019-06-04abs ↗pdf ↗

Paper develops multivariate time series similarity and distance measures.

problem Compensating for misalignments in multivariate time series data.
method Adapted Independent and Dependent DTW strategies to seven elastic similarity and distance measures.
result Each measure achieves highest accuracy on at least one dataset, supporting their value.

Language-based methods improve human similarity approximations without requiring many human judgments.

problem Approximating human similarity judgments using pre-trained deep neural networks (DNNs) is challenging and expensive.
method Developed language-based methods to approximate human similarity judgments, validated with adaptive tag collection pipeline.
result Language-based methods significantly improve performance over DNN-based methods with fewer human judgments.

The problem of hierarchical clustering items from pairwise similarities is found across various scientific disciplines, from biology to networking. Often, applications of clustering techniques are limited by the cost of obtaining similarities between pairs of items. While prior work has been developed to reconstruct cl…

2012-07-19abs ↗pdf ↗

Paper develops a new similarity metric for predicting stock market returns.

problem Predicting stock returns is challenging due to market stochasticity and various influencing factors.
method Case-based reasoning approach using historical pricing data and a novel similarity metric.
result Demonstrates the benefits of the novel similarity metric in predicting stock market returns.

Improves confidence calibration in neural networks by smoothing labels based on class similarity.

problem Improving confidence calibration in deep neural networks for safety-critical applications.
method Proposes a novel label smoothing technique where label values are based on similarities with the reference class, using different similarity measurements.
result Consistently outperforms state-of-the-art calibration techniques on various datasets and network architectures.

A Sim(n-1,1) affine manifold is an affine manifold whose linear holonomy is contained in the similarity lorentzian group but not in the lorentzian group. The class of similarity lorentzian affine manifolds is a small part in the nice class of conformally lorentzian flat manifolds. In this paper we show that a compact S…

2002-07-06abs ↗pdf ↗

In [9] Kaimanovich introduced the concept of augmented tree on the symbolic space of a self-similar set. It is hyperbolic in the sense of Gromov, and it was shown in [13] that under the open set condition, a self-similar set can be identified with the hyperbolic boundary of the tree. In the paper, we investigate in det…

2012-05-16abs ↗pdf ↗

The paper proposes using non-isotropic distances for more accurate trace link recommendation.

problem Time-consuming and error-prone creation and maintenance of trace links.
method Geometric viewpoint on semantic similarity using non-linear similarity measures.
result Non-isotropic distances improve trace link recommendation accuracy.

Modified cosine distance improves similarity performance in data with variance and correlation.

problem Limitations of traditional cosine similarity in random variable spaces with variance and correlation.
method Proposed a variance-adjusted cosine distance metric to overcome limitations of traditional cosine similarity.
result Modified cosine distance shows 100% test accuracy in KNN model on the Wisconsin Breast Cancer Dataset.

If an m+2m+2-manifold MM is locally modeled on $\RR^{m+2}$ with coordinate changes lying in the subgroup $G=\RR^{m+2}\rtimes ({\rO}(m+1,1)\times \RR^+)$ of the affine group ${\rA}(m+2)$, then MM is said to be a \emph{Lorentzian similarity manifold}. A Lorentzian similarity manifold is also a conformally flat Lorentzia…

2011-10-09abs ↗pdf ↗

CLS measures dataset similarity through decision rule performance.

problem Measuring dataset similarity in machine learning, especially for transfer learning and domain adaptation.
method Cross-Learning Score (CLS) measures similarity through bidirectional generalization performance of decision rules, linking to cosine similarity under canonical linear models.
result CLS effectively measures dataset similarity and transferability, validated on synthetic and real-world datasets.

The paper develops a decision support system for hierarchical text classification of conference proceedings.

problem Classifying documents with a fixed hierarchical structure of topics.
method Developed a weighted hierarchical similarity function to calculate topic relevance, using entropy of words to estimate weights.
result The weighted hierarchical similarity function improves ranking accuracy compared to other methods.

We introduce a new discrepancy score between two distributions that gives an indication on their similarity. While much research has been done to determine if two samples come from exactly the same distribution, much less research considered the problem of determining if two finite samples come from similar distributio…

2012-10-15abs ↗pdf ↗

We develop a theory of Nobeling manifolds similar to the theory of Hilbert space manifolds. We show that it reflects the theory of Menger manifolds developed by M. Bestvina and is its counterpart in the realm of complete spaces. In particular, the Nobeling manifold characterization conjecture is proven.

2006-02-27abs ↗pdf ↗

We study the dynamics of correlation and variance in systems under the load of environmental factors. A universal effect in ensembles of similar systems under the load of similar factors is described: in crisis, typically, even before obvious symptoms of crisis appear, correlation increases, and, at the same time, vari…

2009-05-01abs ↗pdf ↗

Matrix factorization is a key component of collaborative filtering-based recommendation systems because it allows us to complete sparse user-by-item ratings matrices under a low-rank assumption that encodes the belief that similar users give similar ratings and that similar items garner similar ratings. This paradigm h…

2016-04-21abs ↗pdf ↗

Similarity algebra extends algebraic structures with quantitative bounds.

problem Exact algebraic structures with strict axioms.
method Framework for approximate algebraic and Lie structures with ε\varepsilon-estimates.
result Similarity structures converge to classical algebraic objects as εightarrow0\varepsilon ightarrow 0.

ROTS improves sentence similarity by incorporating structural information.

problem Measuring sentence similarity with theoretical insights and structural awareness.
method Recursive Optimal Transport (ROT) framework to incorporate structural information.
result ROTS outperforms weakly supervised approaches in sentence similarity tasks.

We propose a novel method of introducing structure into existing machine learning techniques by developing structure-based similarity and distance measures. To learn structural information, low-dimensional structure of the data is captured by solving a non-linear, low-rank representation problem. We show that this low-…

2011-10-26abs ↗pdf ↗

Recently, metric learning and similarity learning have attracted a large amount of interest. Many models and optimisation algorithms have been proposed. However, there is relatively little work on the generalization analysis of such methods. In this paper, we derive novel generalization bounds of metric and similarity …

2012-07-23abs ↗pdf ↗

The development of algorithms for hierarchical clustering has been hampered by a shortage of precise objective functions. To help address this situation, we introduce a simple cost function on hierarchies over a set of points, given pairwise similarities between those points. We show that this criterion behaves sensibl…

2015-10-16abs ↗pdf ↗

There are plenty of problems where the data available is scarce and expensive. We propose a generator of semi-artificial data with similar properties to the original data which enables development and testing of different data mining algorithms and optimization of their parameters. The generated data allow a large scal…

2014-03-28abs ↗pdf ↗

We review some developments on clustering stochastic processes and come with the conclusion that asymptotically consistent clustering algorithms can be obtained when the processes are ergodic and the dissimilarity measure satisfies the triangle inequality. Examples are provided when the processes are distribution ergod…

2019-08-05abs ↗pdf ↗

In a context of document co-clustering, we define a new similarity measure which iteratively computes similarity while combining fuzzy sets in a three-partite graph. The fuzzy triadic similarity (FT-Sim) model can deal with uncertainty offers by the fuzzy sets. Moreover, with the development of the Web and the high ava…

2013-12-21abs ↗pdf ↗

Improves document summarization by combining word embeddings and n-grams.

problem Exact word matching fails to measure semantic similarity between sentences.
method Uses deep embedding features and tf-idf features to improve sentence similarity measure; builds an improved sentence similarity graph; employs a submodular objective function; develops a Transformer-based compression model.
result Outperforms tf-idf based approach and achieves state-of-the-art performance on DUC04 dataset.

Data similarity is a key concept in many data-driven applications. Many algorithms are sensitive to similarity measures. To tackle this fundamental problem, automatically learning of similarity information from data via self-expression has been developed and successfully applied in various models, such as low-rank repr…

2019-03-11abs ↗pdf ↗

A novel approach to analyzing time series generated by complex systems, such as markets, is presented. The basic idea of the approach is the {\it Law of Self-Similar Evolution}, according to which any complex system develops self-similarly. There always exist some internal laws governing the evolution of a system, say …

2001-10-15abs ↗pdf ↗

In this paper, we build up a min-max theory for minimal surfaces using sweepouts of surfaces of genus g2g\geq 2. We develop a direct variational methods similar to the proof of the famous Plateau problem by J. Douglas and T. Rado. As a result, we show that the min-max value for the area functional can be achieved by a …

2011-11-27abs ↗pdf ↗

This paper tackles efficient optimization for nonlinear embeddings in similarity learning.

problem Learning similarity with nonlinear embeddings is challenging due to the large number of pairs.
method Detailed derivations and efficient optimization methods for nonlinear embeddings are developed.
result Efficient optimization methods for nonlinear embeddings are shown to be highly effective.

The special linear groups, the mapping class groups of surfaces, the outer autormorphism groups of free groups appear in numerous domains. Their analogies, developped in particular in K. Vogtmann's work, have been written about a lot. In this report, we concentrate on the contractible spaces on which these groups act i…

2011-10-02abs ↗pdf ↗

Unsupervised text embedding has shown great power in a wide range of NLP tasks. While text embeddings are typically learned in the Euclidean space, directional similarity is often more effective in tasks such as word similarity and document clustering, which creates a gap between the training stage and usage stage of t…

2019-11-04abs ↗pdf ↗