New tools discover latent structure in neural circuits from spike train data.
problem Traditional methods fail to recover neural circuit organization due to noise and temporal dependencies.
method Hierarchical extension of GLM with graph-theoretic priors for latent features and connectivity.
result Reveals latent patterns of neural types and locations from spike trains alone.
Modeling hidden neurons in SNNs using mesoscopic approximations.
problem Underconstrained problem of modeling unobserved neurons in SNNs.
method Coarse-graining and mean-field approximations to derive neuLVM.
result neuLVM can efficiently model large SNNs and recover connectivity parameters.
We develop a fast inference method for non-conjugate Gaussian process models on spike count data.
problem Non-Gaussian spike count data complicates Gaussian Process Factor Analysis.
method We introduce Polynomial Approximate Log-Likelihood (PAL) estimators for non-conjugate GPFA models.
result PAL estimators achieve fast and accurate extraction of latent structure from spike train data.
Gradient descent and SGD can converge to max-margin directions in ReLU models.
problem Understanding the implicit bias of gradient methods in ReLU models.
method Characterization of loss function landscape, analysis of GD and SGD convergence, exploration of multi-neuron network learning.
result Gradient descent and SGD can converge to max-margin directions in ReLU models.
Researchers analyze record statistics in correlated random walks and Lévy flights.
problem Understanding record statistics in correlated time series.
method Review of random walk models and Lévy flights, focusing on number of records and record ages.
result Effects of correlations on record statistics were observed and analyzed.
The study sets performance limits for record linkage using KL divergence.
problem Efficiently merging records in large, noisy databases to remove duplicates.
method Assesses performance bounds using Kullback-Leibler divergence in a Bayesian record linkage framework.
result Provides upper and lower bounds on misclassification probability.
Ancestry improves genealogy search results by ranking diverse record types.
problem Ranking diverse genealogy records equitably from various sources.
method Customized Coordinate Ascent, Stochastic Search, Normalized Cumulative Entropy.
result Demonstrated effectiveness of algorithms in improving relevance and diversity.
We study the statistics of record-breaking events in daily stock prices of 366 stocks from the Standard and Poors 500 stock index. Both the record events in the daily stock prices themselves and the records in the daily returns are discussed. In both cases we try to describe the record statistics of the stock data with…
The study of record statistics of correlated series is gaining momentum. In this work, we study the records statistics of the time series of select stock market data and the geometric random walk, primarily through simulations. We show that the distribution of the age of records is a power law with the exponent α lyi…
Deepr learns features from medical records to predict patient risk.
problem Feature engineering bottleneck in creating predictive systems from medical records.
method Transforms medical records into sequences, uses convolutional neural nets to detect and combine clinical motifs.
result Deepr achieves superior accuracy in predicting patient risk compared to traditional techniques.
New approach protects privacy of deleted records in machine learning.
problem Privacy of deleted records in machine learning models.
method Sound deletion guarantee and noisy gradient descent algorithm.
result Privacy of existing records is necessary for deleted records' privacy.
We present a probabilistic method for linking multiple datafiles. This task is not trivial in the absence of unique identifiers for the individuals recorded. This is a common scenario when linking census data to coverage measurement surveys for census coverage evaluation, and in general when multiple record-systems nee…
Paper proposes a method to locate power grid recordings using ENF sequences.
problem Locating power grid recordings without concurrent power signals.
method Extract ENF sequences from power and audio recordings, develop multi-class SVM model.
result Validation of location authenticity of recordings using ENF sequences.
Paper presents a new method for clustering patient records using tensor decomposition.
problem Clustering high-dimensional binary data, especially in healthcare records.
method Tensor decomposition for an efficient and robust heuristic.
result Clinically meaningful results obtained on two healthcare datasets.
Paper uses ensemblers to predict sepsis early from patient records.
problem Early detection of sepsis in patients.
method Imputation and weak ensembler technique applied to 40k patient records.
result Model achieved 93.45% accuracy and 0.271 utility score.
Extract low-dimensional dynamics from multiple neural recordings.
problem Current methods can't handle dynamics across multiple neural recordings.
method Subspace-identification approach with moment-matching objective and scalable stochastic gradient descent.
result Can identify dynamics and predict correlations even with missing data and small overlap.
Improved neural network detects heart sounds with 87.5% accuracy from noisy recordings.
problem Detecting cardiac abnormalities from noisy heart sound recordings.
method Segmental Convolutional Neural Network (CNN) architecture trained on noisy recordings.
result Best model achieved 87.5% accuracy on PhysioNet/CinC Challenge dataset.
We consider the occurrence of record-breaking events in random walks with asymmetric jump distributions. The statistics of records in symmetric random walks was previously analyzed by Majumdar and Ziff and is well understood. Unlike the case of symmetric jump distributions, in the asymmetric case the statistics of reco…
New method uses surrogate outcomes and single-record data to improve suicide risk modeling.
problem Lack of historical information in single-record patients hinders modeling rare medical events.
method Hybrid framework combining supervised and unsupervised learning to integrate concurrent and single-record data.
result Single-record data and concurrent diagnoses provide valuable information for improving suicide risk modeling.
DNPUs improve neural network performance with high-capacity nanoelectronic nodes.
problem Limited performance of single DNPUs in solving complex classification problems.
method Developed DNPUs as high-capacity neurons and implemented multi-DNPU networks.
result Feed-forward DNPU networks improve single DNPU performance from 77% to 94% test accuracy.
DNI recovers missing brain data from corrupted recordings.
problem Corrupted neural recordings from multielectrode systems.
method Deep Neural Imputation framework using autoencoders.
result DNI recovers both time series and frequency content from corrupted data.
End-to-end DA method for domain-invariant CNNs using parallel audio recordings.
problem Distribution mismatches between training and application data in machine listening.
method Enforcing equal hidden layer representations for domain-parallel samples.
result Learn domain-invariant classifiers without requiring classification labels.
Methodology classifies EEG records using ε-complexity coefficients.
problem Binary classification of multi-channel EEG records for different mental states.
method Extends ε-complexity theory to vector functions, uses coefficients as features. result Accurate classification in four-dimensional space of ε-complexity coefficients. Proposes a novel classifier for probabilistic record linkage.
problem Probabilistic record linkage across databases.
method Graphical model based on mixture of Poisson distributions with latent variables, using gamma priors and supervised labels.
result Classifier works effectively with sparse and streaming data.
Modeling maximum drawdown records in capital markets using PDMP.
problem Capturing the statistical properties of maximum drawdown records in financial markets.
method Piecewise Deterministic Markov Process (PDMP) for modeling, statistical analysis of mean and variance, simulation study, parameter estimation techniques.
result Derivation of statistical results including mean and variance of maximum drawdown records.
Machine learning constructs problem-based medical records from electronic health records.
problem Difficulty in finding relevant medical information for clinical questions.
method Knowledge base completion using machine learning on electronic health records.
result Automatic construction of problem-based groupings of medications, procedures, and lab tests.
Develops an efficient k-means algorithm for clustering incomplete datasets.
problem Clustering datasets with missing values.
method Introduces km-means algorithm that handles incomplete records. result Efficacy demonstrated in various settings and patterns of missing data.
Deep models for financial transactions are vulnerable to adversarial attacks, especially when adding transaction tokens.
problem Vulnerability of deep models to adversarial attacks on financial transaction records.
method Examine adversarial attacks and defenses on transaction records data, considering black-box attacks and adding transaction tokens.
result A few generated transactions can fool a deep-learning model, highlighting the need for robustness improvements.
A machine learning approach to record fusion with high accuracy.
problem Aggregating multiple records corresponding to the same entity.
method Constructing feature vectors from attribute-level, record-level, and database-level signals; using a stagewise additive model to learn a classifier.
result Average precision of ~98% with source information and ~94% without source information across diverse datasets.
Study compares EEG and fMRI systems, finding tradeoffs in artifact removal and classification accuracy.
problem Dealing with artifacts introduced by simultaneous EEG and fMRI recordings.
method Comparison of three MR compatible EEG recording systems, assessing their performance in single-trial EEG classification.
result Tradeoffs across systems, including setup ease and artifact removal methods.
Visual analytics system for comparing medical records using sequence embeddings.
problem Challenges in analyzing medical records due to high dimensionality, irregularity, and sparsity.
method Event and sequence embeddings using autoencoder and self-attention mechanism, with sequence alignment for comparison.
result Demonstrated effectiveness with real-world neonatal ICU dataset.
CorGAN generates synthetic healthcare records while preserving privacy.
problem Generating realistic synthetic healthcare records while maintaining privacy.
method Combining Convolutional Generative Adversarial Networks and Convolutional Autoencoders to capture correlations between medical features.
result CorGAN generates synthetic data with performance similar to real data in various ML settings.
Improves Gaussian process factor models for multi-population recordings.
problem Cubic runtime scaling with trial length and group number limits application to large-scale recordings.
method Two approximate approaches: inducing variables and frequency domain.
result Achieved orders of magnitude speed-up with minimal statistical performance impact.
This study predicts diabetes complications using financial records and neural networks.
problem Managing chronic diseases like diabetes in patients.
method Used financial records from health plans, applied self-attentive recurrent neural networks.
result Successfully predicted diabetes complications with an AUC of 0.81-0.94, 60-240 days ahead.
The task of matching co-referent records is known among other names as rocord linkage. For large record-linkage problems, often there is little or no labeled data available, but unlabeled data shows a reasonable clear structure. For such problems, unsupervised or semi-supervised methods are preferable to supervised met…
XLabel tool reduces medical experts' workload by 40% and explains its decisions.
problem Efficiently labeling large electronic health records.
method Visual-interactive tool using Explainable Boosting Machine (EBM) for classification and explanation.
result EBM achieves high accuracy and explainability, even with mislabeled data.
Bayesian approach for multifile record linkage and duplicate detection.
problem Challenges in merging overlapping datafiles with duplicates.
method Bayesian approach with novel partition representation and loss functions.
result Proposes a flexible prior for partitions and uncertain unresolved portions.
Fetal ECG (FECG) telemonitoring is an important branch in telemedicine. The design of a telemonitoring system via a wireless body-area network with low energy consumption for ambulatory use is highly desirable. As an emerging technique, compressed sensing (CS) shows great promise in compressing/reconstructing data with…
Language models improve clinical prediction models using EHR data.
problem Limited patient data for training clinical prediction models.
method Using patient representation schemes from natural language processing.
result 3.5% mean improvement in AUROC on five prediction tasks.
Biodiversity monitoring using audio recordings is achievable at a truly global scale via large-scale deployment of inexpensive, unattended recording stations or by large-scale crowdsourcing using recording and species recognition on mobile devices. The ability, however, to reliably identify vocalising animal species is…
We study the statistics of records of a one-dimensional random walk of n steps, starting from the origin, and in presence of a constant bias c. At each time-step the walker makes a random jump of length ηdrawn from a continuous distribution f(η) which is symmetric around a constant drift c. We focus in particular on th…
WaveNet reconstructs speech from brain activity, revealing acoustic features.
problem Reconstructing speech from brain activity with limited data.
method WaveNet model applied to STG intracranial recordings.
result WaveNet models reveal phoneme-level acoustic features.
Method reconstructs glacier front trajectories from record moraine data.
problem Understanding past glacier dynamics from limited record data.
method Stochastic generator based on Brownian motion and NBI hyper parameter tuning.
result Reconstructed glacier front trajectories from moraine records.
Scrubbing PHI data from medical records is now efficient and scalable with SpaCy.
problem Efficiency and scalability of de-identification techniques for PHI data.
method Evaluated numerous deep learning techniques including SpaCy for performance and efficiency.
result SpaCy model is both well performing and extremely efficient for PHI data scrubbing.
The study uses demographical data to predict health conditions.
problem Predicting health conditions based on patient demographics and symptoms.
method Analyzing Electronic Health Records (EHR) from Brazil to identify age-related clusters of health conditions.
result Age of patients significantly influences the likelihood of certain health conditions.
Two solutions for multi-modal record linkage using Deep Learning inspired by Visual Question Answering.
problem Matching records from multiple sources representing the same entity.
method Two fusion modules: Recurrent Neural Network + Convolutional Neural Network and Stacked Attention Network. A Siamese Neural Network computes similarity.
result Recurrent Neural Network + Convolutional Neural Network fusion module outperforms a simple model.
While records and order statistics of independent and identically distributed (i.i.d.) random variables X_1, ..., X_N are fully understood, much less is known for strongly correlated random variables, which is often the situation encountered in statistical physics. Recently, it was shown, in a series of works, that one…
We investigate the statistics of records in a random sequence {xB(0)=0,xB(1),⋯,xB(n)=xB(0)=0} of n time steps. The sequence xB(k)'s represents the position at step k of a random walk `bridge' of n steps that starts and ends at the origin. At each step, the increment of the position is a random ju…