CREPE uses deep learning for pitch estimation, outperforming existing methods.
problem Improving pitch estimation accuracy, especially in noisy conditions.
method A deep convolutional neural network trained on time-domain waveforms.
result CREPE achieves state-of-the-art performance in pitch estimation.
Hybrid f0 extraction method for various speech modes with high accuracy.
problem Reliable f0 extraction across different speech modes.
method Ordinal regression CNN and filtering/autocorrelation for pitch estimation.
result Significantly reduces pitch detection error and generalizes to unseen modes.
Deep Autotuner corrects singing pitch using neural networks.
problem Automatic pitch correction for singing performances.
method Neural network model trained on spectrograms of singing and accompaniment.
result Neural network predicts continuous pitch shifts, allowing for improvisation and harmonization.
Model disentangles timbre and pitch for musical instruments.
problem Learning disentangled representations of musical instrument sounds.
method Gaussian mixture variational autoencoders with two separate encoders for timbre and pitch.
result Model successfully disentangles timbre and pitch, enabling controllable synthesis and transfer.
Deep Autotuner corrects singing pitch without scores, using vocal and accompaniment spectral data.
problem Automatic pitch correction without musical scores for singing performances.
method Convolutional Gated Recurrent Unit (CGRU) model trained on karaoke data.
result The model predicts pitch correction from vocal and accompaniment spectral contents, making the voice sound in tune with the accompaniment.
Improved F0 estimation in noisy speech with neural networks.
problem Difficult F0 estimation at low SNRs in unexpected noise.
method Waveform-to-sinusoid regression using RNN trained on supervised data.
result Significant improvement in FPE and GPE rates compared to existing methods.
Improved speech emotion recognition using pitch-synchronous single frequency filtering spectrogram.
problem Uncertainty principles in STFT spectrogram limit time and frequency resolutions.
method Modified SFF spectrogram by averaging amplitudes between GCI locations, named pitch-synchronous SFF spectrogram.
result Improved SER accuracy (63.95% to 70.4%) on IEMOCAP dataset.
In this paper, we investigate the relations between the pitch, the angle of pitch and drall of parallel ruled surface of a closed curve in dual Lorentzian space.
Bayesian Gaussian process model for music audio analysis.
problem Joint estimation of music parameters from audio signals.
method Bayesian Gaussian processes incorporating prior information.
result Improved pitch estimation and missing segment inference.
In this paper, we investigate the relations between the pitch, the angle of pitch and drall of parallel ruled surface of a closed spacelike curve with timelike binormal in dual Lorentzian space.
In this paper, we investigate the relations between the pitch, the angle of pitch and drall of parallel ruled surface of a closed spacelike curve with a spacelike binormal in dual Lorentzian space.
Improved piano transcription by predicting onsets and frames together.
problem Polyphonic piano music transcription accuracy.
method Deep convolutional and recurrent neural network trained to predict pitch onsets and frames.
result Over 100% relative improvement in note F1 score on MAPS dataset.
The study improves pitch detection in polyphonic music by learning harmonic priors.
problem Challenges in transcribing polyphonic music due to overlapping harmonics.
method Introduced Gaussian process priors and used variational Bayes for inference.
result Learning priors that fit the frequency content of sound events improves pitch detection.
Proposes a new voice conversion model that preserves pitch patterns.
problem Preserving pitch patterns while changing speaker identity.
method Variational-autoencoder-based model with an auxiliary network.
result Ensures the conversion result correctly reflects specified F0/timbre information.
Paper proposes a bijective approach for signal/symbol translation using variational auto-encoders.
problem Extracting symbolic information from signals, especially in music, is challenging and non-generic.
method Turned into a density estimation task, using two variational auto-encoders with additive constraint.
result Bijective signal/symbol translation achieved, allowing both signal-to-symbol and symbol-to-signal inference.
Paper improves music transcription models with invariance and data augmentation.
problem Improving accuracy of frame-based music transcription models.
method Translation-invariant network combining filterbank and CNN, trained with pitch-shift augmented data.
result Top-performing model in MIREX evaluation, reducing model complexity and avoiding overfitting.
Music SketchNet generates missing measures in incomplete music pieces, guided by user input.
problem Generating missing measures in incomplete monophonic musical pieces.
method Introducing SketchVAE for factorized representation of rhythm and pitch, and two discriminative architectures for guided music completion.
result Our approach outperforms state-of-the-art models in both objective and subjective evaluations.
Researchers create exact minimal surfaces with helical motifs in biological structures.
problem Analyzing helical motifs in minimal surfaces of biological structures.
method Developed a method to construct exact minimal surfaces with arbitrary helical motifs.
result Exact minimal surfaces with helical motifs can be created and analyzed.
Enhanced Tacotron for Japanese speech synthesis improves naturalness.
problem Challenges in end-to-end Japanese speech synthesis due to pitch accents.
method Extended Tacotron with self-attention to capture pitch accent dependencies.
result Proposed systems show improvements but still lag behind traditional pipeline methods.
HMMs improve music transcription accuracy.
problem Improving automatic transcription of music.
method Employed PLCA for multi-pitch estimation and integrated HMMs for note segmentation and post-processing.
result HMMs enhance transcription accuracy on different instruments.
Improved F0 contour estimation from noisy speech with higher frequency resolution.
problem Estimating the fundamental frequency contour of speech from noisy audio.
method Regression model using recurrent deep neural networks trained on supervised learning.
result Improves F0 contour estimation by more than 25% at SNRs between -10 dB and +10 dB.
EC^2-VAE generates music analogies by disentangling pitch and rhythm representations.
problem Disentangling music representations for generating creative analogies.
method Explicitly-constrained variational autoencoder (EC^2-VAE) for disentangling pitch and rhythm representations.
result EC^2-VAE enables the generation of music analogies by borrowing representations from different pieces.
In this paper, we investigate the asymptotic behavior of regular ends of flat surfaces in the hyperbolic 3-space H^3. Galvez, Martinez and Milan showed that when the singular set does not accumulate at an end, the end is asymptotic to a rotationally symmetric flat surface. As a refinement of their result, we show that …
We present a supervised neural network model for polyphonic piano music transcription. The architecture of the proposed model is analogous to speech recognition systems and comprises an acoustic model and a music language model. The acoustic model is a neural network used for estimating the probabilities of pitches in …
The paper uses neural networks to convert one speaker's voice to another.
problem Identifying speakers uniquely based on their voice.
method Uses convolutional neural networks to manipulate pitch and timbre.
result Preliminary results show encouraging voice conversion.
Paper tackles speaker verification by removing reverberation using deep LSTM networks.
problem Improving speaker verification accuracy in reverberant environments.
method Dual-label deep LSTM networks trained to map reverberant to clean speech features.
result Evaluates performance using EERs, showing improved accuracy.
MIDI-VAE models music dynamics and instrumentation for style transfer.
problem Modeling and transferring musical style across different instruments and dynamics.
method Variational Autoencoder (VAE) for polyphonic music modeling and style transfer.
result MIDI-VAE successfully transfers musical style between different genres and instruments.
TimbreTron transfers musical timbre using CQT and WaveNet.
problem Transfer musical timbre while preserving pitch, rhythm, and loudness.
method Apply image domain style transfer to CQT representation, then generate high-quality waveform with WaveNet.
result TimbreTron recognizably transfers timbre while preserving musical content.
Study shows adding noise to training data improves speech synthesis system's performance under noisy test conditions.
problem Impact of noisy linguistic features on neural network-based speech synthesis systems.
method Comparison of systems using ideal and corrupted linguistic features in training and test sets.
result Adding noise to training data can regularize the model and improve performance under noisy test conditions.
WeSinger improves singing voice synthesis with data augmentation and specialized modules.
problem Improving the accuracy and naturalness of synthesized singing voices.
method Developed a multi-singer Chinese neural singing voice synthesis system with deep bi-directional LSTM, Transformer, LPCNet, and data augmentation.
result WeSinger achieves state-of-the-art performance on the Opencpop corpus.
Geometric characterization of sub-Riemannian geodesics on frame bundles.
problem Characterize sub-Riemannian geodesics on frame bundles of 3-manifolds.
method Lie theoretical description, geometric characterization, complex length spectrum computation.
result Sub-Riemannian metrics on frame bundles of isospectral manifolds are length isospectral.
This paper proposes a method for instrument classification in polyphonic music using monophonic data.
problem Instrument classification in polyphonic music from monophonic data.
method Data augmentation techniques including overlaying audio segments of the same genre, pitch, and tempo synchronization. Convolutional Neural Networks used for classification.
result An ensemble of VGG-like classifiers trained on non-augmented, pitch-synchronized, tempo-synchronized and genre-similar excerpts achieved above 80% LRAP.
For every genus g, we prove that S2×R contains complete, properly embedded, genus-g minimal surfaces whose two ends are asymptotic to helicoids of any prescribed pitch. We also show that as the radius of the S2 tends to infinity, these examples converge smoothly to complete, properly embedded minimal s…
Study of motion control systems on Lie groups with specific geometric constraints.
problem Controlling motion systems on Lie groups with geometric constraints.
method Analysis of control systems on Lie groups, focusing on infinitesimal roto-translations and geodesics.
result Explicit geodesics found for the sub-Riemannian structure on the Lie group.
Machine learning and complexity-entropy methods estimate liquid crystal properties from textures.
problem Extracting physical properties from liquid crystal textures.
method Combining permutation entropy, statistical complexity, and machine learning.
result Significant precision in predicting physical properties of liquid crystals.
SING generates musical notes from instruments in real-time.
problem Efficiently generating high-quality audio from MIDI data.
method Frame-by-frame waveform generation with a single decoder, using a new loss function.
result SING produces significantly improved audio quality compared to state-of-the-art models, with 32x faster training and 2,500x faster inference.
For every genus g, we prove that S^2 x R contains complete, properly embedded, genus-g minimal surfaces whose two ends are asymptotic to helicoids of any prescribed pitch. We also show that as the radius of the S^2 tends to infinity, these examples converge smoothly to complete, properly embedded minimal surfaces in Eu…
DDSP integrates signal processing with deep learning for high-fidelity audio synthesis.
problem Efficiently combining signal processing knowledge with deep learning for audio synthesis.
method Integrates classic signal processing elements with deep learning methods.
result High-fidelity audio synthesis without large models or adversarial losses.
Neural network generates music scores directly from polyphonic audio.
problem Transcribing music scores directly from polyphonic audio.
method Convolutional Recurrent Neural Network (CRNN) with CTC loss function.
result Model can learn to transcribe scores directly from audio signals.
We describe all possible self-similar motions of immersed hypersurfaces in Euclidean space under the mean curvature flow and derive the corresponding hypersurface equations. Then we present a new two-parameter family of immersed helicoidal surfaces that rotate/translate with constant velocity under the flow. We look at…
Improved music transcription accuracy using Particle Filtering for PLCA model.
problem Limited performance of EM-based PLCA models in automatic music transcription.
method Employed Particle Filtering (PF) to overcome EM algorithm's limitations.
result Achieved 61.8% and 59.5% note-level transcription accuracy on two instrument repertoires.
Paper compares XGB and BPNN for music style classification.
problem Efficient music style classification using different methods.
method Feature extraction for timbral texture, rhythmic content, and pitch content; comparative evaluation of XGB and BPNN.
result XGB outperforms BPNN for small datasets in music classification.
Paper introduces a new model for polyphonic music composition.
problem Creating music with multiple interwoven voices.
method Developed a coupled recurrent model using probabilistic factorization and neural network ideas.
result Trained models for single-voice and multi-voice composition on a large dataset.
ConvS2S-VC converts voice characteristics and pitch contour using a fully convolutional seq2seq model.
problem Voice conversion with preservation of pitch contour and duration.
method Fully convolutional seq2seq architecture with conditional batch normalization.
result ConvS2S-VC outperforms baseline methods in sound quality and speaker similarity.
Generative adversarial networks improve speech synthesis from MFCCs.
problem Synthesizing speech from MFCCs, which are typically unusable for synthesis.
method Predict fundamental frequency and voicing from MFCCs, convert spectral envelope to filters, train excitation model, add noise.
result High quality speech can be reconstructed from MFCCs alone.
Paper prunes deep MIR models to ultra-light versions.
problem Massive complexity of deep learning models in MIR.
method Lottery ticket hypothesis-based model pruning.
result Up to 90% of model parameters can be removed without loss of accuracy.
Study on helix curves and their Möbius energy asymptotics.
problem Understanding the asymptotic behavior of Möbius energy for helix curves.
method Investigation of helix curves with fixed radius, focusing on energy decay and blow-up.
result Proven asymptotics for both uncoiling and coiling helix curves, revealing distinct strategies for each.
New algorithm separates vocals from music recordings efficiently.
problem Separate vocal and instrumental parts in music recordings.
method Informed group-sparse representation for linear-time singing voice separation.
result Efficacy confirmed on iKala dataset; music accompaniment follows group-sparse structure.