Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

2579 · Nov 201819922001200920182026
48 results for convincing utterances

New method improves dialogue agents focusing on simple utterances.

problem Dialogue agents often focus on simple utterances and suboptimal policies.
method Tempered Policy Gradient (TPG) methods to improve dialogue performance.
result Significant improvements in dialogue performance, especially in producing convincing utterances.

Modeling long-range context for multi-function utterances in dialogues.

problem Complex dependencies across dialogue turns in long utterances.
method Adapted Convolutional Recurrent Neural Network (CRNN) to model interactions between utterances.
result Significantly outperforms existing work on CDA recognition on a tech forum dataset.

Deep neural network improves speaker verification with short utterances.

problem Challenges in speaker recognition with short utterances.
method Proposes deep neural network mapping methods to improve i-vector performance.
result Deep neural network mapping significantly improves speaker verification performance with short utterances.

Proposes a method to learn speaker embeddings for variable duration utterances.

problem Mismatch between training and testing utterance durations degrades speaker verification performance.
method Sliding window segmentation, LSTM, attentive pooling, segment-level and utterance-level embeddings, similarity loss.
result Significant improvement in robustness for duration variant utterances.

Meta-learning framework improves short utterance speaker recognition.

problem Poor performance of existing models with short utterances.
method Prototypical Networks with support and query sets, enforcing classification against entire training set.
result Significant performance gains on VoxCeleb datasets.

Paper presents a Siamese network for identifying more convincing evidence.

problem Identifying the more convincing argument in discussions.
method Proposes a Siamese neural network architecture for labeling evidence as more convincing.
result The Siamese network outperforms baselines on a new challenging data set.

Improves speaker verification for variable-duration utterances using a feature pyramid module.

problem Improving robustness for variable-duration utterances in speaker verification.
method Integrates a feature pyramid module into multi-scale aggregation to enhance speaker-discriminative information from multiple layers.
result Improves performance for both short and long utterances compared to state-of-the-art approaches.

Graphical lasso models ASR utterance dependencies for consistent WER estimation.

problem Modeling dependent structure among ASR utterances for accurate significance analysis.
method Graphical lasso for dependency modeling, followed by blockwise bootstrap resampling.
result Statistically consistent variance estimator of WER under mild conditions.

Self multi-head attention improves speaker recognition for long utterances.

problem Speaker recognition for long speech segments using Deep Learning.
method Convolutional Neural Network (CNN) for short-term features, self multi-head attention for long-term embeddings.
result Self multi-head attention outperforms other pooling methods by 18% relative EER on VoxCeleb1 dataset.

Improved far-field speaker verification for short utterances in noisy conditions.

problem Challenges in speaker verification on short utterances in uncontrolled noisy environments.
method Used deep neural network architectures (TDNN and ResNet) and experimented with various embedding extractors and training procedures.
result ResNet architectures outperform x-vector approach in speaker verification quality for both long and short utterances.

DGP learns speech recognition by modeling complex relationships between utterances.

problem Modeling complex relationships in speech recognition without relational data.
method Bayesian nonparametric deep learning method (DGP) that generates infinite probabilistic graphs.
result DGP successfully infers relationships among utterances without relational data during training.

A new type of End-to-End system for text-dependent speaker verification is presented in this paper. Previously, using the phonetically discriminative/speaker discriminative DNNs as feature extractors for speaker verification has shown promising results. The extracted frame-level (DNN bottleneck, posterior or d-vector) …

2017-01-03abs ↗pdf ↗

Recurrent sequence generators conditioned on input data through an attention mechanism have recently shown very good performance on a range of tasks in- cluding machine translation, handwriting synthesis and image caption gen- eration. We extend the attention-mechanism with features needed for speech recognition. We sh…

2015-06-24abs ↗pdf ↗

Paper explores unsupervised transfer learning for SLU, improving model performance with unlabeled data.

problem Improving SLU model performance with limited labeled data.
method Uses ELMo embeddings for unsupervised pre-training and ELMo-Light for faster pre-training. Combines unsupervised and supervised transfer techniques.
result Unsupervised pre-training on unlabeled data significantly improves SLU performance, even outperforming conventional supervised transfer.

The paper extracts structured data from physician-patient conversations, reducing clerical burden.

problem Mining insights from physician-patient conversations for electronic health record documentation.
method Created a dataset of transcripts and summaries, extracted noteworthy utterances, and improved model performance.
result Extracting noteworthy utterances significantly boosts model performance for recognizing diagnoses and RoS abnormalities.

Developers of text-to-speech synthesizers (TTS) often make use of human raters to assess the quality of synthesized speech. We demonstrate that we can model human raters' mean opinion scores (MOS) of synthesized speech using a deep recurrent neural network whose inputs consist solely of a raw waveform. Our best models …

2016-11-28abs ↗pdf ↗

The traditional Sznajd model, as well as its Ochrombel simplification for opinion spreading, are applied to marketing with the help of advertising. The larger the lattice is the smaller is the amount of advertising needed to convince the whole market

2002-07-06abs ↗pdf ↗

Proposes a new neural network for text-dependent speaker verification.

problem Improves speaker verification by encoding phrase and speaker information.
method Uses differentiable alignment models to produce supervectors from utterances.
result Achieves competitive performance in text-dependent speaker verification tasks.

Self-supervised method detects replay spoofing using acoustic configurations.

problem Challenges in collecting large-scale datasets for replay spoofing detection.
method Self-supervised pretraining of acoustic configurations using existing datasets.
result The method outperforms baseline by 30% on ASVspoof 2019 physical access dataset.

Blockwise bootstrap improves ASR performance testing for correlated data.

problem Testing reliability of WER improvements between ASR systems.
method Divide evaluation utterances into nonoverlapping blocks and resample these blocks.
result The variance estimator of absolute WER difference is consistent under mild conditions.

A standard recipe for spoken language recognition is to apply a Gaussian back-end to i-vectors. This ignores the uncertainty in the i-vector extraction, which could be important especially for short utterances. A recent paper by Cumani, Plchot and Fer proposes a solution to propagate that uncertainty into the backend. …

2017-09-29abs ↗pdf ↗

Paper introduces methods to automatically generate SOAP notes from patient-physician conversations.

problem Burden of creating digital SOAP notes by physicians.
method Cluster2Sent algorithm for summarizing patient-physician conversations.
result Cluster2Sent algorithm outperforms existing methods by 8 ROUGE-1 points.

Study combines chit-chat and goal-oriented dialogue in fantasy games.

problem Combining naturalistic chit-chat with goal-oriented tasks in fantasy games.
method Trained a goal-oriented model with reinforcement learning against an imitation-learned chit-chat model using two approaches.
result Both models outperform a baseline and can converse naturally to achieve goals.

This paper develops a theory for group Lasso using a concept called strong group sparsity. Our result shows that group Lasso is superior to standard Lasso for strongly group-sparse signals. This provides a convincing theoretical justification for using group sparse regularization when the underlying group structure is …

2009-01-20abs ↗pdf ↗

GE2E-AC improves accent classification by focusing on accent embeddings.

problem Training models to predict accent type can lead to learning irrelevant features.
method GE2E-AC trains models to extract accent embeddings, making them closer for the same accent class.
result GE2E-AC outperforms baseline models trained with conventional loss.

An elementary arbitrage principle and the existence of trends in financial time series, which is based on a theorem published in 1995 by P. Cartier and Y. Perrin, lead to a new understanding of option pricing and dynamic hedging. Intricate problems related to violent behaviors of the underlying, like the existence of j…

2012-06-07abs ↗pdf ↗

Generative x-vectors improve SV performance.

problem Improving text-independent speaker verification.
method Proposes a novel method to combine i-vectors and x-vectors using a transformation model derived from canonical correlation analysis.
result Generative x-vectors provide better performance than baseline systems, especially for long-duration utterances.

The Cartier-Perrin theorem, which was published in 1995 and is expressed in the language of nonstandard analysis, permits, for the first time perhaps, a clear-cut mathematical definition of the volatility of a financial asset. It yields as a byproduct a new understanding of the means of returns, of the beta coefficient…

2011-02-03abs ↗pdf ↗

This paper poses some basic questions about instances (hard to find) of a special problem in 3-manifold topology. "Important though the general concepts and propositions may be with the modern industrious passion for axiomatizing and generalizing has presented us...nevertheless I am convinced that the special problems …

2013-04-18abs ↗pdf ↗

Improves E2E ASR performance on numeric sequences with additional training data and denormalization.

problem Challenges in recognizing numeric sequences out-of-vocabulary in ASR systems.
method Uses text-to-speech for additional numeric training data and a small-footprint neural network for denormalization.
result Reduction of WER by up to a factor of 8 in the longest numeric sequences.