Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

11213242 · Jun 202019922001200920182026
48 results for bilingual dictionary induction

Unsupervised MT struggles with morphologically rich languages.

problem Limitations of unsupervised machine translation on morphologically rich languages.
method Adversarial unsupervised alignment of word embedding spaces for bilingual dictionary induction.
result A simple trick exploiting weak supervision from identical words improves unsupervised bilingual dictionary induction performance.

Geometric approach learns bilingual mappings from monolingual embeddings.

problem Bilingual lexicon induction and cross-lingual word similarity.
method Decouples learning into rotations and a metric, modeled as optimization on Riemannian manifolds.
result Outperforms previous approaches on bilingual lexicon induction and cross-lingual word similarity tasks.

Simple framework decouples word alignment and multilingual embedding mapping.

problem Learning multilingual embeddings without supervision.
method Two-stage approach: 1) unsupervised word alignment, 2) mapping embeddings to shared space.
result Robust performance across various multilingual tasks, including distant languages.

Paper refines cross-lingual word embeddings using Manhattan norm.

problem Sensitivity of 2\ell_{2} norm loss function to outliers in CLWEs.
method Post-processing step using 1\ell_{1} norm to improve CLWEs.
result The 1\ell_{1} refinement substantially outperforms state-of-the-art baselines.

Geometric approach for unsupervised word embedding alignment.

problem Learning alignment between word embeddings of source and target languages.
method Formulates alignment as domain adaptation on the manifold of doubly stochastic matrices, employing Riemannian conjugate gradient algorithm.
result Empirically outperforms state-of-the-art methods on bilingual lexicon induction tasks.

Improved unsupervised word translation using adversarial autoencoder with cycle consistency and input reconstruction.

problem Challenging language pairs and lack of parallel data for unsupervised word translation.
method Adversarial autoencoder with cycle consistency and input reconstruction regularization.
result More stable and better performance than recent approaches.

Our research extends the Bilingual Evaluation Understudy (BLEU) evaluation technique for statistical machine translation to make it more adjustable and robust. We intend to adapt it to resemble human evaluation more. We perform experiments to evaluate the performance of our technique against the primary existing evalua…

2015-09-30abs ↗pdf ↗

Inductive Matrix Completion (IMC) is an important class of matrix completion problems that allows direct inclusion of available features to enhance estimation capabilities. These models have found applications in personalized recommendation systems, multilabel learning, dictionary learning, etc. This paper examines a g…

2016-09-13abs ↗pdf ↗

New method identifies latent variables without strong assumptions.

problem Recovering latent variables from observational data without strong assumptions.
method Diverse dictionary learning, using set-theoretic intersections, complements, and symmetric differences.
result Identifiability of latent variables up to appropriate indeterminacies without strong assumptions.

No free lunch theorems show all algorithms perform equally under uniform distribution.

problem Analyzing scenarios involving non-uniform distributions and comparing algorithms.
method No Free Lunch theorems applied to analyze and compare algorithms without distribution assumptions.
result Anti-cross-validation performs as well as cross-validation under non-uniform distributions.

DeepCAM learns convolutional dictionaries for image processing.

problem Processing high-dimensional signals like images efficiently.
method Introduces a Deep Convolutional Analysis Dictionary Model (DeepCAM) using convolutional dictionaries.
result DeepCAM achieves performance comparable to other methods on single image super-resolution.

Cross-language learning allows us to use training data from one language to build models for a different language. Many approaches to bilingual learning require that we have word-level alignment of sentences from parallel corpora. In this work we explore the use of autoencoder-based methods for cross-language learning …

2014-02-06abs ↗pdf ↗

Bayesian method improves dictionary learning for complex problems.

problem Efficiently identifying relevant dictionary entries for complex inverse problems.
method Bayesian group sparsity coding and deflation steps to compress and identify relevant subdictionaries.
result Significant computational complexity reduction and improved glitch detection in LIGO experiment.

Paper improves dictionary learning by addressing local and global coherence issues.

problem Improving dictionary learning by addressing local and global coherence issues.
method The paper uses the ITKrM algorithm to prove contraction under relaxed conditions and proposes replacing bad dictionaries with carefully designed candidates.
result The adaptive version of ITKrM can recover a generating dictionary from randomly initialized dictionaries of various sizes and learn meaningful dictionaries on image data.

This work learns sparse tensor representations using mixtures of separable dictionaries.

problem Learning sparse representations of tensor data with structured models.
method Proposes and explores learning a mixture of separable dictionaries with sufficient conditions for local identifiability.
result Developed computational algorithms for batch and online learning.

We present a two-stage approach for learning dictionaries for object classification tasks based on the principle of information maximization. The proposed method seeks a dictionary that is compact, discriminative, and generative. In the first stage, dictionary atoms are selected from an initial dictionary by maximizing…

2012-08-17abs ↗pdf ↗

Many techniques in computer vision, machine learning, and statistics rely on the fact that a signal of interest admits a sparse representation over some dictionary. Dictionaries are either available analytically, or can be learned from a suitable training set. While analytic dictionaries permit to capture the global st…

2013-03-21abs ↗pdf ↗

CRsAE auto-encoder recovers convolutional dictionary from noisy signals.

problem Recovering a convolutional dictionary from noisy signals.
method Constrained recurrent sparse auto-encoder (CRsAE) architecture.
result CRsAE successfully recovers the underlying dictionary in the presence of noise.

Paper provides conditions for local recovery of tensor data's Kronecker-structured dictionaries.

problem Local recovery of Kronecker-structured dictionaries for tensor data.
method Derives sufficient conditions for local recovery of coordinate dictionaries.
result Sufficient conditions guarantee recovery of individual coordinate dictionaries up to specified error.

Sparse coding in learned dictionaries has been established as a successful approach for signal denoising, source separation and solving inverse problems in general. A dictionary learning method adapts an initial dictionary to a particular signal class by iteratively computing an approximate factorization of a training …

2012-05-28abs ↗pdf ↗

Study shows unique sharp local minimum in 1\ell_1-minimization for dictionary learning.

problem Global recovery of a dictionary from random linear combinations of atoms.
method Norm condition, explicit bound, perturbation-based test, Block Coordinate Descent algorithm.
result Reference dictionary is the unique sharp local minimum of the 1\ell_1 objective function.

Researchers find optimal dictionaries for minimizing average squared coefficients in random vector representations.

problem Finding optimal dictionaries for minimizing the average squared coefficients in random vector representations.
method Using rank-1 decompositions and majorization theory, the study provides a complete characterization of optimal dictionaries.
result Complete characterization of 2\ell_2-optimal dictionaries with polynomial time algorithms.

In sparse signal representation, the choice of a dictionary often involves a tradeoff between two desirable properties -- the ability to adapt to specific signal data and a fast implementation of the dictionary. To sparsely represent signals residing on weighted graphs, an additional design challenge is to incorporate …

2014-01-05abs ↗pdf ↗

Sparse representations using learned dictionaries are being increasingly used with success in several data processing and machine learning applications. The availability of abundant training data necessitates the development of efficient, robust and provably good dictionary learning algorithms. Algorithmic stability an…

2013-03-03abs ↗pdf ↗

The paper provides guarantees for an alternating minimization algorithm in dictionary learning.

problem Dictionary learning problem of factorizing samples into a basis and sparse vectors.
method Alternating minimization procedure switching between 1\ell_1 minimization and gradient descent.
result Local convergence guarantees for the alternating minimization algorithm under a new matrix infinity norm condition.

The paper tackles dictionary learning with almost sure error constraints.

problem Achieving desirable features in data representation with almost sure error constraints.
method Imposes almost sure recovery constraints and reformulates the problem as a convex-concave min-max problem, solved using gradient descent-ascent.
result Demonstrates the effectiveness of the proposed method in achieving almost sure error constraints in dictionary learning.

We study the Dictionary Learning (aka Sparse Coding) problem of obtaining a sparse representation of data points, by learning \emph{dictionary vectors} upon which the data points can be written as sparse linear combinations. We view this problem from a geometry perspective as the spanning set of a subspace arrangement,…

2014-02-28abs ↗pdf ↗

The paper proposes a method to learn discriminative multilevel dictionaries for supervised image classification.

problem Improving sparse representation for supervised image classification.
method Learning structured multilevel dictionaries with discriminative constraints for each class, using reconstruction errors of image patches.
result Competitive results compared to state-of-the-art methods on texture image classification.

Dictionary learning is a cutting-edge area in imaging processing, that has recently led to state-of-the-art results in many signal processing tasks. The idea is to conduct a linear decomposition of a signal using a few atoms of a learned and usually over-completed dictionary instead of a pre-defined basis. Determining …

2016-05-25abs ↗pdf ↗

Dictionaries are collections of vectors used for representations of random vectors in Euclidean spaces. Recent research on optimal dictionaries is focused on constructing dictionaries that offer sparse representations, i.e., 0\ell_0-optimal representations. Here we consider the problem of finding optimal dictionaries …

2016-03-07abs ↗pdf ↗

This article addresses the issue of representing electroencephalographic (EEG) signals in an efficient way. While classical approaches use a fixed Gabor dictionary to analyze EEG signals, this article proposes a data-driven method to obtain an adapted dictionary. To reach an efficient dictionary learning, appropriate s…

2013-03-04abs ↗pdf ↗

The kernel least-mean-square (KLMS) algorithm is an appealing tool for online identification of nonlinear systems due to its simplicity and robustness. In addition to choosing a reproducing kernel and setting filter parameters, designing a KLMS adaptive filter requires to select a so-called dictionary in order to get a…

2013-10-31abs ↗pdf ↗