Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

14294357 · Oct 201919922001200920182026
48 results for subword units

Improved speech recognition with end-to-end attention models.

problem End-to-end speech recognition with open-vocabulary.
method Sequence-to-sequence attention-based models on subword units, new pretraining scheme, CTC loss function, LSTM language models.
result State-of-the-art word error rates (3.54% and 3.82%) on LibriSpeech test-clean.

New combinatorial framework for geometric realizations of subword complexes.

problem Proving or disproving geometric realizations of subword complexes of Coxeter groups.
method Algebraic combinatorics and discrete geometry framework, parameter matrices.
result Existence of parameter matrices equivalent to realizability of subword complexes as chirotopes.

New techniques reuse subword embeddings in neural models, reducing size and improving performance.

problem Improving performance and reducing model size in subword-aware neural language models.
method Reusing subword embeddings and other weights in multi-layer input embedding models, tying layers consecutively bottom-up.
result Best morpheme-aware model with reused weights outperforms competitive word-level model by a large margin.

The study proves a theorem about subword complexity for free group automorphisms.

problem Analyzing subword complexity for attracting fixed points of automorphisms of free groups.
method Combinatorial arguments and train tracks.
result Subword complexity of attracting fixed points is equivalent to n, n log log n, n log n, or n^2.

Study on representations of four-punctured sphere group in hyperbolic spaces.

problem Understanding representations of the four-punctured sphere group.
method Investigation into simple-stable and Bowditch representations in Gromov-hyperbolic spaces.
result Simple-stable representations and Bowditch representations are equivalent.

Neural model improves text normalization for non-English languages.

problem Improving text normalization in non-English languages with limited data.
method Sequence-to-sequence model with character and word embeddings, using pre-trained word embeddings with subword information.
result Achieved state-of-the-art F1 score on Arabic language correction dataset.

Proposes MorphMine for unsupervised morpheme segmentation to improve word embeddings.

problem Lack of semantic information in word-level analysis for infrequent and out-of-vocabulary words.
method MorphMine applies a parsimony criterion to hierarchically segment words into the fewest number of morphemes.
result MorphMine segments words into human-verified morphemes and improves word embedding quality.

Synthetic noise training improves machine translation robustness to spelling mistakes.

problem Making machine translation robust to spelling mistakes and natural noise.
method Training on synthetic noise to improve robustness to natural noise.
result Training on synthetic noise improves robustness to natural noise without diminishing performance on clean text.

Probabilistic FastText captures multiple word senses and sub-word structures.

problem Capturing multiple word senses and sub-word structures in word embeddings.
method Probabilistic FastText uses Gaussian mixture densities to represent words, sharing statistical strength across sub-word structures and capturing different word senses.
result Probabilistic FastText outperforms existing models on word-similarity benchmarks and discerning different meanings.

SAFER method certifies robustness to word substitutions without model structure.

problem Certified robustness against synonymous word substitutions in NLP models.
method Randomized smoothing with stochastic ensemble of randomized inputs.
result Significantly outperforms state-of-the-art methods for certified robustness.

The study proves biharmonic unit sections on 2-tori are always harmonic and exists in each homotopy class.

problem Characterizing biharmonic unit vector fields and sections on 2-tori.
method Analyzing variational problems for unit vector fields under conformal metrics, proving properties through homotopy classes.
result Biharmonic unit sections on 2-tori are always harmonic and exist in each homotopy class.

In this paper we propose and investigate a novel nonlinear unit, called LpL_p unit, for deep neural networks. The proposed LpL_p unit receives signals from several projections of a subset of units in the layer below and computes a normalized LpL_p norm. We notice two interesting interpretations of the LpL_p unit. First…

2013-11-07abs ↗pdf ↗

Randomly chosen primary hidden units and derived secondary units reduce neural network complexity.

problem Large number of hidden units in neural networks.
method Introducing primary and secondary hidden units with random weights for primary units and derived weights for secondary units.
result Significant reduction in the number of hidden units without compromising accuracy.

Study examines how business units can benefit from group cohesion under regulatory constraints.

problem Regulatory constraints limit business units' ability to form a single cohesive group.
method Defined and analyzed cohesive risk measures to minimize capital costs.
result Cohesive risk measures allow groups to achieve minimal capital costs without altering individual liabilities.

Study examines dependence properties of Bayesian neural network units in finite-width networks.

problem Understanding dependence properties of hidden units in practical finite-width Bayesian neural networks.
method Theoretical analysis and empirical evaluation of depth and width impacts.
result Hidden units in finite-width Bayesian neural networks are dependent, contrary to the infinite-width limit assumption.

We study deep Bayesian neural networks with Gaussian priors, revealing heavy-tailed unit activations.

problem Characterizing regularization effects in deep Bayesian neural networks.
method Investigation of deep Bayesian neural networks with Gaussian weight priors and ReLU-like nonlinearities.
result The prior distribution on units becomes increasingly heavy-tailed with depth, influencing activation patterns.

Wasserstein t-SNE embeds hierarchical datasets considering within-unit distributions.

problem Exploring hierarchical datasets where units are compared based on means of sample distributions.
method Uses Wasserstein distance metric for 2D embeddings of units, approximating Gaussian distributions for efficiency.
result Demonstrates effective embedding of hierarchical datasets, uncovering meaningful structure.

Study on hidden units in finite Bayesian neural networks and their tail properties.

problem Understanding the behavior of hidden units in finite Bayesian neural networks.
method Introduced a generalized Weibull-tail property to describe hidden units tails.
result Unit priors become heavier-tailed going deeper, providing insights into finite Bayesian neural networks.

We present a new equation with respect to a unit vector field on Riemannian manifold MnM^n such that its solution defines a totally geodesic submanifold in the unit tangent bundle with Sasaki metric and apply it to some classes of unit vector fields. We introduce a class of covariantly normal unit vector fields and pro…

2005-09-30abs ↗pdf ↗

Neural Power Unit (NPU) learns arbitrary power functions on real numbers.

problem Neural Networks struggle with generalizing beyond seen data and arithmetic operations.
method Introduces Neural Power Unit (NPU) that operates on real numbers and learns arbitrary power functions.
result NPU outperforms competitors in accuracy and sparsity on arithmetic datasets and discovers governing equations from data.

ReLU units can 'die' in neural networks, causing slower convergence.

problem ReLU units sometimes produce near-zero outputs during training.
method Simulation and statistical analysis of a simplified ReLU unit model.
result Activation probability decreases as training progresses, leading to slower convergence.

Bayesian SHMM discovers acoustic units from unlabeled speech.

problem Discovering language-specific acoustic units from unlabeled speech.
method Bayesian Subspace Hidden Markov Model (SHMM) trained on labeled data to find new acoustic units on target language.
result Significantly outperforms previous HMM-based systems and compares favorably with Variational Auto Encoder-HMM.

We construct homotopically non-trivial maps from the unit m-sphere to the unit (m-1)-sphere with arbitrarily small k-dilation for each k greater than (m + 1)/2. We prove that homotopically non-trivial maps from the unit m-sphere to the unit (m-1)-sphere cannot have arbitrarily small k-dilation for k less than or equal …

2012-11-05abs ↗pdf ↗

We present a probabilistic variant of the recently introduced maxout unit. The success of deep neural networks utilizing maxout can partly be attributed to favorable performance under dropout, when compared to rectified linear units. It however also depends on the fact that each maxout unit performs a pooling operation…

2013-12-20abs ↗pdf ↗