Proposes a speaker-independent GlotNet vocoder using WaveNet for speech generation.
problem Lack of efficient multi-speaker WaveNet models with limited resources.
method Uses source-filter model of speech production to train a WaveNet for glottal excitation.
result Proposed GlotNet vocoder performs favorably to direct WaveNet vocoder in speech quality.
Paper aims to find joint representation between vocal tract geometry and speech sound acoustics.
problem Finding a joint latent representation between articulatory and acoustic domains for vowel sounds.
method Invertible neural network models, convolutional autoencoder, normalizing flows, semi-supervised learning.
result Satisfactory performance in articulatory-to-acoustic and acoustic-to-articulatory mapping.
Breath sounds can identify speakers with high accuracy.
problem Identifying speakers from breath sounds.
method Examined breath sounds during continuous speech, focusing on inhalation phase.
result Breath sounds carry unique speaker-specific information, enabling accurate speaker identification.
End-to-end speech recognition system trained on GPUs and CPUs.
problem Building state-of-the-art speech recognition systems.
method Utilizes CPUs and GPUs for training, data augmentation, and neural network updates. Uses vocal tract length perturbation and acoustic simulator for data augmentation. Employed Horovod allreduce for training.
result Achieved 7.92% WER on proprietary English Bixby open domain test set using a Bidirectional Full Attention (BFA) model.
Geometric framework for aligning fiber tracts across subjects.
problem Challenges in finding direct tract correspondence across multiple individuals.
method Geometric framework using intrinsic mean and deformation fields, parallel transport for registration.
result Bundle alignment results on 43 healthy adult subjects.
Novel 3D U-Net method for fast, reproducible white matter tract segmentation.
problem Challenges in fast and consistent white matter tract segmentation from diffusion tensor MRI.
method Convolutional neural network (3D U-Net) trained on a large DTI dataset.
result Reproducibility and accuracy of tract-specific diffusion measures.
New method clusters infant vocalizations using topological data.
problem Clustering infant vocalizations for developmental analysis.
method Topologically augmented signal representation with Dirichlet process mixture model.
result 8 clusters of vocalizations identified in the first 12 months of life.
Framework converts singer identity and vocal technique from non-parallel corpora.
problem Converts singer identity and vocal technique from non-parallel corpora.
method Uses variational autoencoders with separate encoders for singer identity and vocal technique.
result Successfully disentangles and converts singer identity and vocal technique.
A fusion approach combines audio and video features for emotion recognition.
problem Continuous emotion recognition using both visual and auditory modalities.
method Pre-trained CNN features from video frames and minimalistic auditory descriptors. Fusion at feature or prediction level. SVR for prediction.
result Improves CCCs of 0.749 and 0.565 for arousal and valence respectively.
Understanding how housing values evolve over time is important to policy makers, consumers and real estate professionals. Existing methods for constructing housing indices are computed at a coarse spatial granularity, such as metropolitan regions, which can mask or distort price dynamics apparent in local markets, such…
Cheap model diagnoses vocal disorders accurately.
problem Diagnosing vocal disorders without expensive equipment.
method Used Mel-Cepstrum vectors and Support Vector Machine.
result Accurately diagnosed three vocal disorders.
Deep learning model classifies gastrointestinal diseases with high accuracy.
problem Disease detection in the gastrointestinal tract.
method Global features and deep neural networks.
result 95.80% accuracy, 95.87% precision, 95.80% F1-score.
Paper analyzes human vocal sentiment using various techniques.
problem Improving accuracy in emotion-level classification of human vocal expressions.
method Conventional vocal feature extraction, deep-learning approaches, context-level analysis, hyperparameter sweeps, data augmentation.
result Improved performance in emotion-level classification.
Study improves machine learning models for GI tract disease detection using comprehensive evaluations and cross-dataset testing.
problem Incomplete or incorrect evaluation of machine learning models for GI tract diseases.
method Comprehensive evaluations of five machine learning models using Global Features and Deep Neural Networks, introducing performance hexagons and cross-dataset testing.
result Demonstrates the need for more sophisticated performance metrics and evaluation methods to build generalizable models.
Study examines equity in post-Snow Uri recovery, finds disparities.
problem Disproportionate impacts on vulnerable populations during recovery.
method County and census tract level data analysis, satellite imagery, statistical procedures.
result Negative associations between non-Hispanic whites and outages, positive associations with certain demographic variables.
Deep learning model estimates multiple f0s, melodies, vocals, and bass lines from music.
problem Estimating f0s and other musical elements from polyphonic music.
method Multitask deep learning architecture trained on a large dataset.
result Multitask model outperforms single-task models.
Deep Autotuner corrects singing pitch without scores, using vocal and accompaniment spectral data.
problem Automatic pitch correction without musical scores for singing performances.
method Convolutional Gated Recurrent Unit (CGRU) model trained on karaoke data.
result The model predicts pitch correction from vocal and accompaniment spectral contents, making the voice sound in tune with the accompaniment.
Proposes deep learning method for GCI detection from pathological speech.
problem Detecting glottal closure instants (GCI) in pathological acoustic speech.
method Convolutional neural network with fused deep acoustic speech and linear prediction residual features.
result Significantly better than state-of-the-art methods in GCI detection.
A new model cleans vocal note event annotations in music.
problem Erroneous labels in music datasets.
method Contrastive learning to automatically create local deformations of likely correct labels.
result Transcription model accuracy improves with the proposed strategy.
Wave-U-Net with MHE regularization improves singing voice separation.
problem Singing voice separation from mixed music recordings.
method Wave-U-Net architecture with MHE regularization applied to 1D filters.
result Adding MHE regularization to the loss function consistently improves singing voice separation.
Hybrid CNN improves segmentation and registration of white matter tracts.
problem Accurate analysis of longitudinal brain imaging data.
method A hybrid CNN integrating segmentation and registration into a single procedure.
result Hybrid CNN outperforms multistage pipelines in segmentation accuracy, consistency, and speed.
New algorithm separates vocals from music recordings efficiently.
problem Separate vocal and instrumental parts in music recordings.
method Informed group-sparse representation for linear-time singing voice separation.
result Efficacy confirmed on iKala dataset; music accompaniment follows group-sparse structure.
DC-SIS selects features faster than mRMR for Parkinson's vocal diagnosis.
problem Feature selection for Parkinson's disease vocal data.
method DC-SIS (Distance Correlation Sure Independence Screening) using distance correlation measure.
result 90 times faster feature selection with similar accuracy.
ConvNet classifies whale vocalizations and ambient noise in acoustic recordings.
problem Automated detection and classification of marine mammal vocalizations in acoustic recordings.
method Convolutional Neural Network with a novel acoustic representation.
result Classifier accurately detects and classifies whale vocalizations and ambient noise.
SCM-GAN converts any song to sound like a different singer.
problem Converting songs to sound like a different artist.
method Transfer learning and GANs to separate vocals and instruments, then convert and merge.
result SCM-GAN improves song conversion metrics by 35% GV and 13% MS.
This paper addresses converting speech to EGG signals without hardware, improving accuracy.
problem Estimating EGG signals from speech without hardware.
method Optimization of evidence lower bound with KL-divergence minimization.
result The method generates EGG signals that agree with gold standard and outperforms state-of-the-art.
New method separates music vocals from accompaniment without labeled data.
problem Separating music sources without isolated recordings.
method Bootstrapping deep model using primitive auditory cues.
result Trained deep model separates vocals from accompaniment in unlabeled music.
Paper proposes efficient multivariate spatial Fay-Herriot models using variational autoencoders.
problem Estimating population characteristics in small areas with limited data.
method Integrates multivariate spatial Fay-Herriot model with variational autoencoders to leverage spatial structure efficiently.
result Significant computational efficiency improvements for high-dimensional datasets.
A new neural network separates vocals from music accompaniment.
problem Separating vocals from music accompaniment in recordings.
method Self-attention convolutional neural network (CNN) with densely-connected blocks.
result 19.5% relative improvement in vocals separation.
New method provides reliable probabilistic bounds for VUR detection.
problem Detect VUR in children without radiation exposure.
method Machine learning with probabilistic bounds for conditional probability.
result Guaranteed bounds contain well-calibrated probabilities.
A faster Bayesian method for estimating spatial count data models.
problem Bayesian estimation of spatial count data models is computationally expensive and slow.
method Derive a Variational Bayes (VB) method for posterior inference in negative binomial models with spatial dependence.
result The VB method is up to 50 times faster than MCMC and offers similar accuracy.
Hierarchical CNNs improve diagnosis of GI diseases from histopathological images.
problem Diagnosing GI diseases from histopathological images is challenging due to heterogeneity and shared features.
method Embedded a class hierarchy into a VGGNet to address the hierarchical structure of GI diseases.
result The hierarchical model achieved better results than a flat model for multi-category diagnosis of GI disorders.
We extend the well-known Denjoy-Ahlfors theorem on the number of different asymptotic tracts of holomorphic functions to subharmonic functions on arbitrary Riemannian manifolds. We obtain some new versions of the Liouville theorem for $\p$-harmonic functions without requiring the geodesic completeness requirement of a …
Continuous speech recognition from brain activity without vocalization.
problem Recognizing silent speech from EEG signals.
method Implemented a CTC ASR model using EEG signals.
result Demonstrated feasibility of EEG for continuous silent speech recognition.
Efficiently generates and selects explanations for neural networks using GANs and FID.
problem Manual selection of hyper-parameters for generating interpretable neural network explanations is slow and requires qualitative evaluation.
method Proposes a novel metric using Fréchet Inception Distance (FID) and a GAN-based method for efficient search and realistic output generation.
result Successfully selects hyper-parameters leading to interpretable examples, avoiding manual evaluation.
In this work we show that the systems of balance equations (balance systems) of continuum thermodynamics occupy a natural place in the variational bicomplex formalism. We apply the vertical homotopy decomposition to get a local splitting (in a convenient domain) of a general balance system as the sum of a Lagrangian pa…
Study uses machine learning to identify IBD biomarkers from gut microbiota.
problem Identifying biomarkers for Inflammatory Bowel Disease (IBD) from gut microbiota.
method Ensemble feature selection methods (CMIM, FCBF, mRMR, XGBoost) applied to IBD-associated metagenomics dataset.
result XGBoost minimizes microbiota used for IBD diagnosis, improving classification accuracy.
Jukebox generates high-fidelity songs with singing in raw audio.
problem Generating music with singing in raw audio.
method Multi-scale VQ-VAE for compression, autoregressive Transformers for modeling.
result Generates high-fidelity and diverse songs with coherence up to multiple minutes.
The human auditory system is able to distinguish the vocal source of thousands of speakers, yet not much is known about what features the auditory system uses to do this. Fourier Transforms are capable of capturing the pitch and harmonic structure of the speaker but this alone proves insufficient at identifying speaker…
Bird sound data collected with unattended microphones for automatic surveys, or mobile devices for citizen science, typically contain multiple simultaneously vocalizing birds of different species. However, few works have considered the multi-label structure in birdsong. We propose to use an ensemble of classifier chain…
Paper presents unsupervised learning for visual representations using patches from unlabelled videos.
problem Learning visual representations without labeled data.
method Trains a model for foreground and background classification using patches extracted from unlabelled videos.
result Model achieves 45.3 mAP, close to best unsupervised learning techniques.
Improved music source separation using spectrogram feature loss.
problem Music source separation quality improvement.
method Added a high-level feature loss term from spectrograms using a VGG net to a deep learning model.
result Improvement in separation quality of drums and vocals from songs.
Paper proves spectral filters can be transferred between graphs.
problem Proving spectral filters can be transferred between graphs.
method Introducing the Cayley smoothness space and proving filters in this space are linearly stable.
result Graph spectral filters are transferable if they are in the Cayley smoothness space.
Diffusion magnetic resonance imaging (dMRI) and tractography provide means to study the anatomical structures within the white matter of the brain. When studying tractography data across subjects, it is usually necessary to align, i.e. to register, tractographies together. This registration step is most often performed…
This work prunes CNN filters based on their functionality, not just size.
problem Redundant filters in CNNs waste computation resources.
method Functionality-oriented filter pruning method.
result Pruning based on functionality optimizes computation and interprets filter importance.
A new SOHP filter improves trend estimation in economic time series.
problem Improving trend estimation in nonlinear economic time series.
method Recursive application of one-sided HP filter on updated cyclical components, combined with an incremental HP filtering algorithm.
result Better performance of SOHP filter compared to other HP-type filters on real economic data.
Deep density methods improve filtering in high-dimensional systems.
problem Nonlinear filtering in high-dimensional systems.
method Two deep density methods based on Feynman-Kac formulas and neural networks.
result Logarithmic deep backward stochastic differential equation filter outperforms classical methods in high dimensions.
Enhances geodesic fiber tracking in white matter using modified metrics and tensor data.
problem Improving the accuracy and robustness of geodesic fiber tracking in white matter.
method Modification of geodesic ray-tracing method using rescaled metrics and fourth-order tensor data.
result More satisfactory results in the construction of white matter tracts as geodesics.