Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

0.3%0.6%0.8%1.1% · Sep 201019922001200920182026
48 results for pitch accent

Enhanced Tacotron for Japanese speech synthesis improves naturalness.

problem Challenges in end-to-end Japanese speech synthesis due to pitch accents.
method Extended Tacotron with self-attention to capture pitch accent dependencies.
result Proposed systems show improvements but still lag behind traditional pipeline methods.

GE2E-AC improves accent classification by focusing on accent embeddings.

problem Training models to predict accent type can lead to learning irrelevant features.
method GE2E-AC trains models to extract accent embeddings, making them closer for the same accent class.
result GE2E-AC outperforms baseline models trained with conventional loss.

This paper describes the data collection effort that is part of the project Sprekend Nederland (The Netherlands Talking), and discusses its potential use in Automatic Accent Location. We define Automatic Accent Location as the task to describe the accent of a speaker in terms of the location of the speaker and its hist…

2016-02-08abs ↗pdf ↗

State-of-the-art automatic speech recognition (ASR) systems struggle with the lack of data for rare accents. For sufficiently large datasets, neural engines tend to outshine statistical models in most natural language processing problems. However, a speech accent remains a challenge for both approaches. Phonologists ma…

2018-07-09abs ↗pdf ↗

ConvS2S-VC converts voice characteristics and pitch contour using a fully convolutional seq2seq model.

problem Voice conversion with preservation of pitch contour and duration.
method Fully convolutional seq2seq architecture with conditional batch normalization.
result ConvS2S-VC outperforms baseline methods in sound quality and speaker similarity.

Model disentangles timbre and pitch for musical instruments.

problem Learning disentangled representations of musical instrument sounds.
method Gaussian mixture variational autoencoders with two separate encoders for timbre and pitch.
result Model successfully disentangles timbre and pitch, enabling controllable synthesis and transfer.

Deep Autotuner corrects singing pitch without scores, using vocal and accompaniment spectral data.

problem Automatic pitch correction without musical scores for singing performances.
method Convolutional Gated Recurrent Unit (CGRU) model trained on karaoke data.
result The model predicts pitch correction from vocal and accompaniment spectral contents, making the voice sound in tune with the accompaniment.

Improved speech emotion recognition using pitch-synchronous single frequency filtering spectrogram.

problem Uncertainty principles in STFT spectrogram limit time and frequency resolutions.
method Modified SFF spectrogram by averaging amplitudes between GCI locations, named pitch-synchronous SFF spectrogram.
result Improved SER accuracy (63.95% to 70.4%) on IEMOCAP dataset.

The task of estimating the fundamental frequency of a monophonic sound recording, also known as pitch tracking, is fundamental to audio processing with multiple applications in speech processing and music information retrieval. To date, the best performing techniques, such as the pYIN algorithm, are based on a combinat…

2018-02-17abs ↗pdf ↗

Improved LID for multilingual speakers using context-aware models.

problem Low accuracy for languages spoken by multilingual speakers, especially with accented speech.
method Coarser-grained acoustic model and integration with interaction context signals.
result Average 97% accuracy across all language combinations, 60% improvement in worst-case accuracy.

Music SketchNet generates missing measures in incomplete music pieces, guided by user input.

problem Generating missing measures in incomplete monophonic musical pieces.
method Introducing SketchVAE for factorized representation of rhythm and pitch, and two discriminative architectures for guided music completion.
result Our approach outperforms state-of-the-art models in both objective and subjective evaluations.

We advance the state of the art in polyphonic piano music transcription by using a deep convolutional and recurrent neural network which is trained to jointly predict onsets and frames. Our model predicts pitch onset events and then uses those predictions to condition framewise pitch predictions. During inference, we r…

2017-10-30abs ↗pdf ↗

The aim of this article is to present the category of bounded Frechet manifolds in respect to which we will review the geometry of Frechet manifolds with a stronger accent on its metric aspect. An inverse function theorem in the sense of Nash and Moser in this category is proved, and some applications to Riemannian geo…

2006-12-14abs ↗pdf ↗

Researchers create exact minimal surfaces with helical motifs in biological structures.

problem Analyzing helical motifs in minimal surfaces of biological structures.
method Developed a method to construct exact minimal surfaces with arbitrary helical motifs.
result Exact minimal surfaces with helical motifs can be created and analyzed.

Unified analysis of efficient local training methods for distributed variational inequalities.

problem Efficient distributed/federated learning for variational inequality problems.
method Unified convergence analysis of communication-efficient local training methods.
result First local gradient descent-accent algorithms with improved communication complexity.

EC^2-VAE generates music analogies by disentangling pitch and rhythm representations.

problem Disentangling music representations for generating creative analogies.
method Explicitly-constrained variational autoencoder (EC^2-VAE) for disentangling pitch and rhythm representations.
result EC^2-VAE enables the generation of music analogies by borrowing representations from different pieces.

In this paper, we investigate the asymptotic behavior of regular ends of flat surfaces in the hyperbolic 3-space H^3. Galvez, Martinez and Milan showed that when the singular set does not accumulate at an end, the end is asymptotic to a rotationally symmetric flat surface. As a refinement of their result, we show that …

2007-08-02abs ↗pdf ↗

The human auditory system is able to distinguish the vocal source of thousands of speakers, yet not much is known about what features the auditory system uses to do this. Fourier Transforms are capable of capturing the pitch and harmonic structure of the speaker but this alone proves insufficient at identifying speaker…

2016-10-27abs ↗pdf ↗

Paper proposes a bijective approach for signal/symbol translation using variational auto-encoders.

problem Extracting symbolic information from signals, especially in music, is challenging and non-generic.
method Turned into a density estimation task, using two variational auto-encoders with additive constraint.
result Bijective signal/symbol translation achieved, allowing both signal-to-symbol and symbol-to-signal inference.

MIDI-VAE models music dynamics and instrumentation for style transfer.

problem Modeling and transferring musical style across different instruments and dynamics.
method Variational Autoencoder (VAE) for polyphonic music modeling and style transfer.
result MIDI-VAE successfully transfers musical style between different genres and instruments.

WeSinger improves singing voice synthesis with data augmentation and specialized modules.

problem Improving the accuracy and naturalness of synthesized singing voices.
method Developed a multi-singer Chinese neural singing voice synthesis system with deep bi-directional LSTM, Transformer, LPCNet, and data augmentation.
result WeSinger achieves state-of-the-art performance on the Opencpop corpus.

This study compares PRNG and QRNG in machine learning models, revealing significant differences in performance.

problem Implications of PRNG and QRNG on machine learning model performances.
method Used CPU and QPU to generate random numbers for various machine learning techniques.
result Quantum Random Number Generators (QRNG) outperform Pseudo Random Number Generators (PRNG) in certain tasks.

Geometric characterization of sub-Riemannian geodesics on frame bundles.

problem Characterize sub-Riemannian geodesics on frame bundles of 3-manifolds.
method Lie theoretical description, geometric characterization, complex length spectrum computation.
result Sub-Riemannian metrics on frame bundles of isospectral manifolds are length isospectral.

This paper proposes a method for instrument classification in polyphonic music using monophonic data.

problem Instrument classification in polyphonic music from monophonic data.
method Data augmentation techniques including overlaying audio segments of the same genre, pitch, and tempo synchronization. Convolutional Neural Networks used for classification.
result An ensemble of VGG-like classifiers trained on non-augmented, pitch-synchronized, tempo-synchronized and genre-similar excerpts achieved above 80% LRAP.

We present a supervised neural network model for polyphonic piano music transcription. The architecture of the proposed model is analogous to speech recognition systems and comprises an acoustic model and a music language model. The acoustic model is a neural network used for estimating the probabilities of pitches in …

2015-08-07abs ↗pdf ↗

For every genus gg, we prove that S2×RS^2 \times R contains complete, properly embedded, genus-gg minimal surfaces whose two ends are asymptotic to helicoids of any prescribed pitch. We also show that as the radius of the S2S^2 tends to infinity, these examples converge smoothly to complete, properly embedded minimal s…

2015-08-01abs ↗pdf ↗

Study of motion control systems on Lie groups with specific geometric constraints.

problem Controlling motion systems on Lie groups with geometric constraints.
method Analysis of control systems on Lie groups, focusing on infinitesimal roto-translations and geodesics.
result Explicit geodesics found for the sub-Riemannian structure on the Lie group.

SING generates musical notes from instruments in real-time.

problem Efficiently generating high-quality audio from MIDI data.
method Frame-by-frame waveform generation with a single decoder, using a new loss function.
result SING produces significantly improved audio quality compared to state-of-the-art models, with 32x faster training and 2,500x faster inference.

For every genus g, we prove that S^2 x R contains complete, properly embedded, genus-g minimal surfaces whose two ends are asymptotic to helicoids of any prescribed pitch. We also show that as the radius of the S^2 tends to infinity, these examples converge smoothly to complete, properly embedded minimal surfaces in Eu…

2013-04-22abs ↗pdf ↗

In this paper, we present a reverberation removal approach for speaker verification, utilizing dual-label deep neural networks (DNNs). The networks perform feature mapping between the spectral features of reverberant and clean speech. Long short term memory recurrent neural networks (LSTMs) are trained to map corrupted…

2018-09-08abs ↗pdf ↗

We describe all possible self-similar motions of immersed hypersurfaces in Euclidean space under the mean curvature flow and derive the corresponding hypersurface equations. Then we present a new two-parameter family of immersed helicoidal surfaces that rotate/translate with constant velocity under the flow. We look at…

2011-06-22abs ↗pdf ↗