Compressed LLM embeddings improve noisy regression tasks without overfitting.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Kernel-spectral embedding learns low-dim. structures from noisy data.
Kernel method embeds noisy datasets, capturing shared structures.
Measures time-delay embedding for noisy, sparse data.
This work proposes a hybrid method for error detection in noisy Knowledge Graphs.
Develops a method to estimate treatment effects using noisy proxies over time.
This research tackles unsupervised topic extraction in noisy social media data.
A novel method extracts topological features from word embeddings for text classification.
We propose -graph embedding for robustly learning feature vectors from data vectors and noisy link weights. A newly introduced empirical moment -score reduces the influence of contamination and robustly measures the difference between the underlying correct expected weights of links and the specified generative m…
This paper proposes an out-of-sample extension framework for a global manifold learning algorithm (Isomap) that uses temporal information in out-of-sample points in order to make the embedding more robust to noise and artifacts. Given a set of noise-free training data and its embedding, the proposed framework extends t…
Paper proposes a matrix optimization model for reliable Euclidean embedding from noisy data.
Improves scalability and robustness of dynamic graph clustering.
We present a novel view of nonlinear manifold learning using derivative-free optimization techniques. Specifically, we propose an extension of the classical multi-dimensional scaling (MDS) method, where instead of performing gradient descent, we sample and evaluate possible "moves" in a sphere of fixed radius for each …
IterefinE combines KG refinement with embeddings to improve KG quality.
To investigate objects without a describable notion of distance, one can gather ordinal information by asking triplet comparisons of the form "Is object closer to or is closer to ?" In order to learn from such data, the objects are typically embedded in a Euclidean space while satisfying as many triplet …
(Bolukbasi et al., 2016) demonstrated that pretrained word embeddings can inherit gender bias from the data they were trained on. We investigate how this bias affects downstream classification tasks, using the case study of occupation classification (De-Arteaga et al.,2019). We show that traditional techniques for debi…
New method detects corporate fraud in noisy financial networks.
Change detection in dynamic networks is an important problem in many areas, such as fraud detection, cyber intrusion detection and health care monitoring. It is a challenging problem because it involves a time sequence of graphs, each of which is usually very large and sparse with heterogeneous vertex degrees, resultin…
We present a novel methodology able to distinguish meaningful level shifts from typical signal fluctuations. A two-stage regularization filtering can accurately identify the location of the significant level-shifts with an efficient parameter-free algorithm. The developed methodology demands low computational effort an…
Author2Vec generates user embeddings from social media data.
We find the minimax rate of convergence in Hausdorff distance for estimating a manifold M of dimension d embedded in R^D given a noisy sample from the manifold. We assume that the manifold satisfies a smoothness condition and that the noise distribution has compact support. We show that the optimal rate of convergence …
With the increasing popularity of video sharing websites such as YouTube and Facebook, multimodal sentiment analysis has received increasing attention from the scientific community. Contrary to previous works in multimodal sentiment analysis which focus on holistic information in speech segments such as bag of words re…
Low dimensional embeddings that capture the main variations of interest in collections of data are important for many applications. One way to construct these embeddings is to acquire estimates of similarity from the crowd. However, similarity is a multi-dimensional concept that varies from individual to individual. Ex…
SymNoise improves language model fine-tuning by 6.7% over NEFTune, using symmetric noise.
Improved far-field speaker verification for short utterances in noisy conditions.
Paper introduces a novel reward function for noisy financial markets using imitation learning.
Existing dimensionality reduction methods are adept at revealing hidden underlying manifolds arising from high-dimensional data and thereby producing a low-dimensional representation. However, the smoothness of the manifolds produced by classic techniques over sparse and noisy data is not guaranteed. In fact, the embed…
In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker. We achieve this by training two separate neural networks: (1) A speaker recognition network that produces speaker-discriminative embeddings; (2) …
In this work we approach the task of learning multilingual word representations in an offline manner by fitting a generative latent variable model to a multilingual dictionary. We model equivalent words in different languages as different views of the same word generated by a common latent variable representing their l…
Dropout improves MIL performance on noisy WSI classification.
A new method improves generative models by learning lower-dimensional representations.
Beta-SOD detects and corrects noisy object re-identification using cosine similarity and Beta mixtures.
Improves domain adaptation by aligning source and target distributions and mitigating noisy labels.
A new algorithm optimizes unknown functions with noisy data and unmatched features.
Unified framework for hyperbolic embeddings from mixed data types.
New framework discovers PDEs from sparse, noisy data.
Entity linking is the task of mapping potentially ambiguous terms in text to their constituent entities in a knowledge base like Wikipedia. This is useful for organizing content, extracting structured data from textual documents, and in machine learning relevance applications like semantic search, knowledge graph const…
Many data-rich industries are interested in the efficient discovery and modelling of structures underlying large data sets, as it allows for the fast triage and dimension reduction of large volumes of data embedded in high dimensional spaces. The modelling of these underlying structures is also beneficial for the creat…
Learning workable representations of dynamical systems is becoming an increasingly important problem in a number of application areas. By leveraging recent work connecting deep neural networks to systems of differential equations, we propose \emph{variational integrator networks}, a class of neural network architecture…
MPVAE learns latent embeddings and label correlations for multi-label classification.
We consider the problem of learning the nearest neighbor graph of a dataset of n items. The metric is unknown, but we can query an oracle to obtain a noisy estimate of the distance between any pair of items. This framework applies to problem domains where one wants to learn people's preferences from responses commonly …
A new topology design improves zero-shot classification performance in contrastive learning.
Study methods to recover unknown processes in PDEs from data.
Manifold learning and dimensionality reduction techniques are ubiquitous in science and engineering, but can be computationally expensive procedures when applied to large data sets or when similarities are expensive to compute. To date, little work has been done to investigate the tradeoff between computational resourc…
Recent research on network embedding in hyperbolic space have proven successful in several applications. However, nodes in real world networks tend to interact through several distinct channels. Simple aggregation or ignorance of this multiplexity will lead to misleading results. On the other hand, there exists redunda…
Develops statistical confidence sets for multidimensional scaling.
The performance of a Part-of-speech (POS) tagger is highly dependent on the domain ofthe processed text, and for many domains there is no or only very little training data available. This work addresses the problem of POS tagging noisy user-generated text using a neural network. We propose an architecture that trains a…
We consider the problem of finding a target object using pairwise comparisons, by asking an oracle questions of the form \emph{"Which object from the pair is more similar to ?"}. Objects live in a space of latent features, from which the oracle generates noisy answers. First, we consider the {\em non-bli…