Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

3537071,0601,413 · Jun 202019922001200920182026
48 results for word piece models

Improved speech recognition model reduces error rate from 9.2% to 5.6%.

problem Improving speech recognition accuracy for voice search tasks.
method Structural and optimization improvements to Listen, Attend, and Spell (LAS) model.
result Significant reduction in word error rate (WER) from 9.2% to 5.6% on voice search task.

Paper proposes a deep learning method to measure domain similarity in Persian texts.

problem Measuring the similarity between different domains of Persian text descriptions.
method Built a dataset of paired texts, used word embeddings and deep neural networks to score similarity, trained on GPU.
result Best model achieved an F1 score of 0.9865.

The paper proposes a novel approach to music analysis using text mining techniques.

problem Analyzing musical documents using traditional text mining methods.
method Developed a Naive Dictionary of 'muselets' (musical words) of uniform length.
result Demonstrated reasonable topic modeling and pattern recognition results with a simplified dictionary.

Study growth patterns in random networks using i.i.d. perturbations.

problem Understanding the growth of affine regions in random piecewise-linear networks.
method Analyzes a random compositional model with i.i.d. perturbations of the tent map, proving submultiplicative pressure and using finite-state defect process for upper-tail lower bounds.
result Proves the existence of a submultiplicative pressure for \(N_n\) and gives exponential upper bounds for \(n^{-1}\log N_n\).

The paper proposes a privacy-preserving method for text data using Hyperbolic space.

problem Preserving user privacy in text data while maintaining utility for machine learning.
method Word representations in Hyperbolic space to provide privacy, sampling from a probability distribution.
result Demonstrates significant privacy guarantees (20x greater) compared to Euclidean space.

A Bloom filter approach combined with Transformer models improves accuracy for machine learning tasks on opaque IDs.

problem Improving accuracy for machine learning tasks on opaque IDs with large vocabulary sizes.
method Applying hash functions to map opaque IDs to multiple hash tokens, similar to a Bloom filter, and using a multi-layer Transformer to process these digests.
result Models outperform those without hashing and sampled softmax, achieving high accuracy with a smaller computational budget.

We introduce causal pieces to improve spiking neural networks.

problem Improving the expressiveness and trainability of spiking neural networks.
method We decompose the input domain of SNNs into causal regions, proving that the number of these regions is a measure of SNNs' approximation capabilities.
result Parameter initialisations yielding a high number of causal pieces correlate with SNN training success.

Bardo Composer generates tabletop RPG music based on player speech.

problem Creating immersive background music for tabletop RPGs.
method Speech recognition, emotion classification, and music generation using a novel beam search algorithm.
result Generated music pieces can be accurately identified by human subjects as conveying the intended emotion.

New method counts boundary pieces in ReLU classifiers for better complexity measure.

problem Current classification complexity measures are misleading and ineffective.
method Developed a novel method using tropical geometry to count exact boundary pieces.
result Boundary piece count is negatively correlated with robustness.

New method for reworking music pieces using conditional autoregressive modeling.

problem Reimagining existing musical compositions while maintaining structure.
method Conditional autoregressive pipeline for efficient music recomposition.
result Diverse and structured generation of new content conditioned on chord sequence annotations.

Estimates change-points and graph structures in a time-varying Ising model.

problem Detecting and understanding changes in a time-varying Ising model.
method Maximizing a penalized conditional log-likelihood to estimate neighborhood of each node, enforcing sparsity and piece-wise constant graph structures.
result First change-points consistency theorems for unknown number of change-points in time-varying Ising model.

Bordered Floer homology assigns invariants to 3-manifolds with boundary, such that the Heegaard Floer homology of a closed 3-manifold, split into two pieces, can be recovered as a tensor product of the bordered invariants of the pieces. We construct cornered Floer homology invariants of 3-manifolds with codimension-2 c…

2013-08-31abs ↗pdf ↗

Machine learning identifies Shakespeare and Fletcher's contributions to Henry VIII.

problem Determining the relative contributions of Shakespeare and Fletcher in Henry VIII.
method Combined analysis of vocabulary and versification with machine learning techniques.
result Supports canonical division and new modifications of Henry VIII's authorship.

The Heegaard genus g of an irreducible closed orientable 3-manifold puts a limit on the number and complexity of the pieces that arise in the Jaco-Shalen-Johannson decomposition of the manifold by its canonical tori. For example, if p of the complementary components are not Seifert fibered, then p < g. This result gene…

1998-02-20abs ↗pdf ↗

A topologically minimal surface may be isotoped into a normal form with respect to a fixed triangulation. If the intersection with each tetrahedron is simply connected, then the pieces of this normal form are triangles, quadrilaterals, and helicoids. Helical pieces can have any number of positive or negative twists. We…

2015-02-19abs ↗pdf ↗

PyChEst detects changes in non-stationary time series without distributional assumptions.

problem Detecting changes in non-stationary time series data.
method Nonparametric algorithms for consistent detection of multiple changepoints in piece-wise stationary processes.
result PyChEst consistently detects changes without distributional assumptions.

Probabilistic FastText captures multiple word senses and sub-word structures.

problem Capturing multiple word senses and sub-word structures in word embeddings.
method Probabilistic FastText uses Gaussian mixture densities to represent words, sharing statistical strength across sub-word structures and capturing different word senses.
result Probabilistic FastText outperforms existing models on word-similarity benchmarks and discerning different meanings.

This paper proposes a method to approximate non-Gaussian likelihoods in Gaussian Processes.

problem Approximating non-Gaussian likelihoods in Gaussian Processes.
method Proposes a piece-wise constant approximation for the inverse-link function.
result Yields a closed form solution for the SVGP lower bound.

XOFM explains attribute effects in ordinal regression using piece-wise linear functions.

problem Lack of detailed attribute contributions in existing ordinal regression models.
method XOFM uses piece-wise linear functions to approximate attribute contributions and introduces ordinal transformation.
result XOFM provides superior explainability and state-of-the-art prediction accuracy.

End-to-end ASR model combines word and character representation for improved performance.

problem Difficulty in training with word-level supervision due to sparsity of examples.
method Multi-task learning framework combining word and character representations.
result Improved word-error rate (WER) by interpolating between word-level and character-level models.

MIDI-VAE models music dynamics and instrumentation for style transfer.

problem Modeling and transferring musical style across different instruments and dynamics.
method Variational Autoencoder (VAE) for polyphonic music modeling and style transfer.
result MIDI-VAE successfully transfers musical style between different genres and instruments.

Develops a robust model for skewed and heavy-tailed data in periodontal studies.

problem Skewed and heavy-tailed data in periodontal pocket depth measurements.
method Flexible two-piece scale Student-t error distribution and deep neural network with monotonicity constraints.
result Robust mode-based estimation resistant to outliers with clinical interpretability.

We define and give explicit construction of the universal tree-graded space with a given collection of pieces. We apply that to proving uniqueness of asymptotic cones of relatively hyperbolic groups whose peripheral subgroups have unique asymptotic cones. Modulo the Continuum Hypothesis, we show that if an asymptotic c…

2010-10-18abs ↗pdf ↗

In this paper we construct quasiconformal embeddings from Y-pieces that contain a short boundary geodesic into degenerate ones. These results are used in a companion paper to study the Jacobian tori of Riemann surfaces that contain small simple closed geodesics.

2013-11-04abs ↗pdf ↗

A new method selects anchor words for better topic discovery in text corpora.

problem Selecting anchor words for improved topic modeling in text corpora.
method Proposes a new greedy method to find a minimum edge-weight anchor clique in a word similarity graph.
result The proposed method outperforms existing methods on topic quality and is faster.

This paper closes the accuracy gap between A2W models and sub-word models using English conversational speech data.

problem Closing the accuracy gap between direct acoustics-to-word models and sub-word models.
method Training an A2W model with orders of magnitude more data, optimizing model initialization, training data order, and regularization.
result Achieved word error rates of 8.8%/13.9% on Hub5-2000 Switchboard/CallHome test sets.

Sparse activations in neural models correlate with frequent words, suggesting sparsity is natural.

problem Interpretability and resource efficiency in neural language models.
method Used the Taxi-Euclidean norm to measure sparsity and analyzed gradients and activations of frequent words.
result Frequent input words are associated with sparse activations, while frequent target words are associated with dispersed activations.

Bayesian method detects change points and clusters in piece-wise constant signals.

problem Detecting change points and clustering in piece-wise constant signals.
method Nonparametric penalized least square model selection on partitions of design points, with an efficient algorithm.
result Oracle inequality and adaptive upper bound on expected square risk of the estimator.

Characterizes fundamental groups of disjointly tree-graded spaces.

problem Understanding fundamental groups of complex geometric structures.
method Defines and analyzes disjointly tree-graded spaces, characterizing their fundamental groups.
result Fundamental groups of disjointly tree-graded spaces embed into inverse limits of free products of fundamental groups of pieces.

The abstract explains how word and relation representations capture semantic meaning.

problem Understanding how word and relation representations capture semantic meaning.
method Theoretical justification and extension of geometric relationships between word embeddings and knowledge graph representations.
result The geometric relationships between word embeddings correspond to semantic relations between words and entities in knowledge graphs.

With a simple architecture and the ability to learn meaningful word embeddings efficiently from texts containing billions of words, word2vec remains one of the most popular neural language models used today. However, as only a single embedding is learned for every word in the vocabulary, the model fails to optimally re…

2017-06-08abs ↗pdf ↗

Paper analyzes word embedding composition using tensor decomposition.

problem Given vector representations of two words, compute a vector for the entire phrase.
method Generative model with low rank Tucker decomposition of word embedding correlations.
result Word embeddings and a core tensor can be derived from the Tucker decomposition.

Most existing word embedding approaches do not distinguish the same words in different contexts, therefore ignoring their contextual meanings. As a result, the learned embeddings of these words are usually a mixture of multiple meanings. In this paper, we acknowledge multiple identities of the same word in different co…

2016-11-29abs ↗pdf ↗