Sequence-based model predicts protein-protein interactions with high accuracy.
problem Predicting protein-protein interactions for alternative treatment options.
method Sequence clustering, discrete cosine transform, supervised machine learning, SVM with RBF.
result Mesh model achieved an average AUC of 0.84.
Predicts cryptocurrency pump probability using sequence-based neural networks.
problem Detecting pump-and-dump schemes in cryptocurrency markets.
method Developed a sequence-based neural network (SNN) that encodes historical P&D events into sequences for prediction.
result SNN improves prediction accuracy by leveraging positional attention to extract useful information.
PANDA predicts protein binding affinity changes from sequences, outperforming existing methods.
problem Accurately predicting changes in protein binding affinity due to mutations.
method Sequence-based machine learning approach using protein sequence information.
result PANDA achieves higher Pearson correlation coefficients than existing methods.
A new method for fast graph embedding using diffusion graphs.
problem Efficiently generating graph embeddings for large networks.
method Diffusion graphs for rapid vertex sequence generation.
result Improved accuracy and performance with higher edge density.
This paper discusses about an R package that implements the Pattern Sequence based Forecasting (PSF) algorithm, which was developed for univariate time series forecasting. This algorithm has been successfully applied to many different fields. The PSF algorithm consists of two major parts: clustering and prediction. The…
Proposes a validator model to ensure valid sequences in deep generative models.
problem Invalid sequences hinder the utility of deep generative models for discrete spaces.
method Introduces a deep recurrent validator model to estimate sequence validity and constrain model output.
result Demonstrates improved validity in Python code and molecular structures.
New method uses Cantor embeddings and Wasserstein distances to analyze predictive states in time series data.
problem Analyzing predictive states in stochastic processes using time series data.
method Wasserstein distances for detecting predictive equivalences in symbolic data, using Cantor embeddings for finite-dimensional representation.
result Exploratory analysis of temporal structure in various processes reveals insights.
Study predicts price predictability in ultra-high frequency financial data using entropy tests.
problem Tackles predictability of ultra-high frequency financial data.
method Develops statistical tests based on Shannon entropy and Kullback-Leibler divergence to analyze predictability.
result Degree of randomness increases with aggregation level in transaction time.
Neural networks predict TED Talk ratings from transcripts, removing bias.
problem Predicting public speaking performance from speech transcripts.
method Causal diagram modeling, word sequence and dependency tree based neural networks.
result Average F-score of 0.77, significantly outperforming baseline methods.
Seq2Seq models speed up epidemic model predictions.
problem Complex epidemic models are computationally expensive.
method Used deep seq2seq models as surrogates for complex models.
result Surrogates predict scenarios up to several thousand times faster.
Our model improves sequence prediction accuracy and diversity.
problem Predicting future sequences with uncertainty and diversity.
method Gaussian Latent Variable model with 'Best of Many' sample objective.
result Empirically outperforms prior work on diverse tasks.
A CNN-based model improves stock price prediction accuracy.
problem Overfitting in image-based stock prediction models.
method SMSFR-CNN combining CNN and image features.
result SMSFR-CNN achieves high predictive accuracy on A-share stocks.
We consider the setting of sequential prediction of arbitrary sequences based on specialized experts. We first provide a review of the relevant literature and present two theoretical contributions: a general analysis of the specialist aggregation rule of Freund et al. (1997) and an adaptation of fixed-share rules of He…
Deep neural networks predict B-cell epitopes for SARS-CoV and SARS-CoV-2.
problem Accurate prediction of B-cell epitopes for vaccine design.
method Deep neural network model with regularization techniques and key features analysis.
result Overall accuracy of 82% in predicting COVID-19 cases.
Enhances solar flare prediction with advanced preprocessing and contrastive learning.
problem Accurate prediction of solar flares to mitigate risks to astronauts and equipment.
method Advanced data preprocessing pipeline and contrastive learning with GRU regression model.
result Exceptional True Skill Statistic (TSS) scores, surpassing previous methods.
Two deep learning models predict protein-protein interactions with high accuracy.
problem Overfitting and information leak in deep learning models for PPI prediction.
method Carefully designed deep learning models, strict conditions for training and testing, and methodology to avoid information leak.
result Best model predicts more than 78% of human PPI with strong confidence.
ISAAC audits deep models for drug-target interactions, revealing structural differences.
problem Deep models for DTI often use irrelevant features, making them hard to evaluate.
method ISAAC uses intervention-based structural auditing to evaluate model sensitivity.
result ISAAC reveals significant structural differences in DTI models' reasoning.
Szabó recently introduced a combinatorially-defined spectral sequence in Khovanov homology. After reviewing its construction and explaining our methodology for computing it, we present results of computations of the spectral sequence. Based on these computations, we make a number of conjectures concerning the structure…
Bayes-assisted confidence sequences improve efficiency for bounded means.
problem Efficient uncertainty quantification for bounded IID means without parametric assumptions.
method Bayesian working predictive model selects adaptive martingale updates maximizing predictive log-growth.
result Asymptotically log-optimal performance with informative priors reducing width and sampling effort.
EcoCast predicts biodiversity risks using satellite data and citizen science records.
problem Unprecedented shifts in species distributions due to climate change and habitat loss.
method Spatio-temporal model using sequence-based transformers and continual learning.
result Promising improvements in forecasting bird species distributions compared to Random Forest.
Prototype Matching Network (PMN) improves genomic TFBS prediction.
problem Predicting Transcription Factor Binding Sites (TFBSs) with hundreds of TFs as labels.
method Prototype Matching Network (PMN) that learns motif-like features and TF-TF interactions.
result PMN significantly outperforms baselines on a large TFBS dataset.
Framework for causal discovery using multi-modal data.
problem Failure of representation learning in causal tasks.
method Statistical and computational framework combining representation learning and causal inference.
result Effective use of observational and perturbational data for causal discovery.
ChemBoost predicts protein-ligand binding affinity using SMILES syntax.
problem Predicting high affinity drug-target interactions from sequence similarity alone.
method ChemBoost uses SMILES syntax to represent ligands as documents and proteins as sequences or ligand-centric features. It learns chemical word embeddings and predicts affinities using eXtreme Gradient Boosting.
result ChemBoost outperforms state-of-the-art systems in predicting protein-ligand affinities.
DeepAffinity predicts compound-protein affinity from sequences, outperforming existing methods.
problem Lack of methods to predict compound-protein affinity from sequences alone.
method Unified RNN/GCNN-CNN model that unifies recurrent and convolutional neural networks.
result Model outperforms conventional options in predicting affinities with high accuracy.
Machine learning identifies math sequences based on empirical laws.
problem Identifying interesting mathematical structures.
method Extract features from integer sequences using Benford's and Taylor's laws; experiment with classifiers.
result Machine learning can identify various mathematical properties in sequences.
CoSE models complex drawings by treating strokes as a collection that can be composed.
problem Modeling complex free-form structures like diagrams.
method Generative model using autoencoder to project strokes into a fixed latent space, relational model operates in latent space.
result Model captures both individual strokes and their compositional structure.
Novel parallel GNN predicts protein-ligand interactions with high accuracy.
problem Accurate prediction of protein-ligand interactions for drug design.
method Parallel Graph Neural Networks (GNN) integrating 3D structural data.
result GNN achieves high accuracy in predicting binary interactions and activity.
New method maps protein sequences to embeddings encoding structural information.
problem Inferring structural properties from amino acid sequences when structures are unknown.
method Representation learning using bidirectional LSTM models with structural similarity and residue contact maps.
result Trained embeddings improve structural similarity prediction and transfer to other tasks.
New method classifies protein structures using network features.
problem Efficiently predicting protein function from structural data.
method Modelled protein structures as PSNs, used graphlets and deep learning for features.
result Proposed methods outperform existing PSC approaches in accuracy.
A new method predicts protein functions using variable-length sequences.
problem Computational methods for protein function prediction are slow and inaccurate for long sequences.
method Two feature sets: single fixed-sized segments and multi-sized segments, using bi-directional LSTM. Combined with MLDA features.
result Significant improvement in accuracy for long protein sequences.
Study local Weyl law on hyperbolic surfaces, identifying geodesic loops.
problem Understanding the variance of a local Weyl law on hyperbolic surfaces.
method Explicit integration of test functions, stationary phase arguments, and geometric analysis of geodesic loops.
result Identifies length-minimizing geodesic loops and sequences, proving they are simple.
Nowadays, hyperspectral image classification widely copes with spatial information to improve accuracy. One of the most popular way to integrate such information is to extract hierarchical features from a multiscale segmentation. In the classification context, the extracted features are commonly concatenated into a lon…
Hidden semi-Markov models (HSMMs) are latent variable models which allow latent state persistence and can be viewed as a generalization of the popular hidden Markov models (HMMs). In this paper, we introduce a novel spectral algorithm to perform inference in HSMMs. Unlike expectation maximization (EM), our approach cor…
Graph-structured data appears frequently in domains including chemistry, natural language semantics, social networks, and knowledge bases. In this work, we study feature learning techniques for graph-structured inputs. Our starting point is previous work on Graph Neural Networks (Scarselli et al., 2009), which we modif…
Deep neural networks improve recommendation accuracy in marketplaces.
problem Measuring and optimizing recommender performance in marketplaces.
method Hybrid item representation models, sequence-based models, and multi-armed bandit models.
result Promising deep neural network recommenders are currently in production at FINN.no.
Novel ligand-based method improves protein representation performance.
problem Improving protein representation for bioinformatics tasks.
method Proposes SMILESVec method to represent ligands and compute protein similarity.
result Ligand-based protein representation performs as well as sequence-based methods.
Deep neural network translates math formula images to LaTeX sequences.
problem Translating math formula images to LaTeX sequences accurately and efficiently.
method Encoder-decoder architecture with CNN and LSTM, sequence-level training with policy gradient.
result State-of-the-art performance on sequence-based and image-based evaluation metrics.
Modeling a temporal process as if it is Markovian assumes the present encodes all of the process's history. When this occurs, the present captures all of the dependency between past and future. We recently showed that if one randomly samples in the space of structured processes, this is almost never the case. So, how d…
Paper proves convergence of MDL to Einstein-Hilbert with boundary term.
problem Proving convergence of discrete MDL to continuous Einstein-Hilbert action.
method Proves \(Γ\)-convergence using diffeomorphism-natural discrete MDL-type functional.
result Identifies Carathéodory densities and obtains \(\liminf/\limsup\) bounds.
Enhanced text-to-speech synthesizes expressive speech from a single example.
problem Creating a new expressive speech style from a single example of speech.
method Combines VAE and Normalizing Flows to improve disentanglement and naturalness.
result Reduces KL-divergence by 22% and improves perceptual metrics.
Adversarial learning for mixture Hawkes processes improves performance.
problem Learning mixture models of Hawkes processes from event sequences.
method Iterative self-paced learning with adversarial self-paced mechanism.
result The proposed method outperforms traditional methods consistently.
A new model aligns sequences using DPMM, outperforming GP-LVM.
problem Aligning high-dimensional time-warped sequences without supervision.
method Dirichlet Process Mixture Model (DPMM) with Gaussian Processes (GPs).
result DPMM achieves competitive results compared to GP-LVM on synthetic and real-world data.
In a 1967 paper, Banchoff stated that a certain type of polyhedral curvature, that applies to all finite polyhedra, was zero at all vertices of an odd-dimensional polyhedral manifold; one then obtains an elementary proof that odd-dimensional manifolds have zero Euler characteristic. In a previous paper, the author defi…
Deviance-style normalization for sparse, jointly overdispersed count matrices
problem Jointly overdispersed count matrices
method Dirichlet-multinomial deviance residualization
result Preserves exact sparsity, evaluates in constant time, recovers multinomial residual
The paper proposes a generic learning method for complex structures.
problem Learning algorithms for complex structures are inefficient and require significant effort.
method Mapping any complex structure onto a generic form (serialization) and applying sequence-based density estimators.
result The method is competitive or better than specialized algorithms for given structures and provides protection from overfitting.
This paper studies when particle filtering is efficient for planning in partially observed systems.
problem The efficiency of particle filtering for planning in partially observed linear dynamical systems.
method Coupling of ideal and approximate sequences to bound particle complexity.
result Polynomially many particles suffice for stable systems to approximate optimal planning.
CR-AIS improves AIS efficiency by constant rate annealing.
problem Efficiently sample from intractable distributions.
method Constant rate annealing schedule for AIS.
result CR-AIS outperforms existing Adaptive AIS methods.
Enzyme sequences and structures are routinely used in the biological sciences as queries to search for functionally related enzymes in online databases. To this end, one usually departs from some notion of similarity, comparing two enzymes by looking for correspondences in their sequences, structures or surfaces. For a…