Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

126251377502 · Jun 202019922001200920182026
48 results for Sequence-based Prediction

Sequence-based model predicts protein-protein interactions with high accuracy.

problem Predicting protein-protein interactions for alternative treatment options.
method Sequence clustering, discrete cosine transform, supervised machine learning, SVM with RBF.
result Mesh model achieved an average AUC of 0.84.

Predicts cryptocurrency pump probability using sequence-based neural networks.

problem Detecting pump-and-dump schemes in cryptocurrency markets.
method Developed a sequence-based neural network (SNN) that encodes historical P&D events into sequences for prediction.
result SNN improves prediction accuracy by leveraging positional attention to extract useful information.

PANDA predicts protein binding affinity changes from sequences, outperforming existing methods.

problem Accurately predicting changes in protein binding affinity due to mutations.
method Sequence-based machine learning approach using protein sequence information.
result PANDA achieves higher Pearson correlation coefficients than existing methods.

Proposes a validator model to ensure valid sequences in deep generative models.

problem Invalid sequences hinder the utility of deep generative models for discrete spaces.
method Introduces a deep recurrent validator model to estimate sequence validity and constrain model output.
result Demonstrates improved validity in Python code and molecular structures.

New method uses Cantor embeddings and Wasserstein distances to analyze predictive states in time series data.

problem Analyzing predictive states in stochastic processes using time series data.
method Wasserstein distances for detecting predictive equivalences in symbolic data, using Cantor embeddings for finite-dimensional representation.
result Exploratory analysis of temporal structure in various processes reveals insights.

Study predicts price predictability in ultra-high frequency financial data using entropy tests.

problem Tackles predictability of ultra-high frequency financial data.
method Develops statistical tests based on Shannon entropy and Kullback-Leibler divergence to analyze predictability.
result Degree of randomness increases with aggregation level in transaction time.

Neural networks predict TED Talk ratings from transcripts, removing bias.

problem Predicting public speaking performance from speech transcripts.
method Causal diagram modeling, word sequence and dependency tree based neural networks.
result Average F-score of 0.77, significantly outperforming baseline methods.

We consider the setting of sequential prediction of arbitrary sequences based on specialized experts. We first provide a review of the relevant literature and present two theoretical contributions: a general analysis of the specialist aggregation rule of Freund et al. (1997) and an adaptation of fixed-share rules of He…

2012-07-09abs ↗pdf ↗

Deep neural networks predict B-cell epitopes for SARS-CoV and SARS-CoV-2.

problem Accurate prediction of B-cell epitopes for vaccine design.
method Deep neural network model with regularization techniques and key features analysis.
result Overall accuracy of 82% in predicting COVID-19 cases.

Enhances solar flare prediction with advanced preprocessing and contrastive learning.

problem Accurate prediction of solar flares to mitigate risks to astronauts and equipment.
method Advanced data preprocessing pipeline and contrastive learning with GRU regression model.
result Exceptional True Skill Statistic (TSS) scores, surpassing previous methods.

Two deep learning models predict protein-protein interactions with high accuracy.

problem Overfitting and information leak in deep learning models for PPI prediction.
method Carefully designed deep learning models, strict conditions for training and testing, and methodology to avoid information leak.
result Best model predicts more than 78% of human PPI with strong confidence.

ISAAC audits deep models for drug-target interactions, revealing structural differences.

problem Deep models for DTI often use irrelevant features, making them hard to evaluate.
method ISAAC uses intervention-based structural auditing to evaluate model sensitivity.
result ISAAC reveals significant structural differences in DTI models' reasoning.

Bayes-assisted confidence sequences improve efficiency for bounded means.

problem Efficient uncertainty quantification for bounded IID means without parametric assumptions.
method Bayesian working predictive model selects adaptive martingale updates maximizing predictive log-growth.
result Asymptotically log-optimal performance with informative priors reducing width and sampling effort.

EcoCast predicts biodiversity risks using satellite data and citizen science records.

problem Unprecedented shifts in species distributions due to climate change and habitat loss.
method Spatio-temporal model using sequence-based transformers and continual learning.
result Promising improvements in forecasting bird species distributions compared to Random Forest.

Prototype Matching Network (PMN) improves genomic TFBS prediction.

problem Predicting Transcription Factor Binding Sites (TFBSs) with hundreds of TFs as labels.
method Prototype Matching Network (PMN) that learns motif-like features and TF-TF interactions.
result PMN significantly outperforms baselines on a large TFBS dataset.

ChemBoost predicts protein-ligand binding affinity using SMILES syntax.

problem Predicting high affinity drug-target interactions from sequence similarity alone.
method ChemBoost uses SMILES syntax to represent ligands as documents and proteins as sequences or ligand-centric features. It learns chemical word embeddings and predicts affinities using eXtreme Gradient Boosting.
result ChemBoost outperforms state-of-the-art systems in predicting protein-ligand affinities.

DeepAffinity predicts compound-protein affinity from sequences, outperforming existing methods.

problem Lack of methods to predict compound-protein affinity from sequences alone.
method Unified RNN/GCNN-CNN model that unifies recurrent and convolutional neural networks.
result Model outperforms conventional options in predicting affinities with high accuracy.

Machine learning identifies math sequences based on empirical laws.

problem Identifying interesting mathematical structures.
method Extract features from integer sequences using Benford's and Taylor's laws; experiment with classifiers.
result Machine learning can identify various mathematical properties in sequences.

Novel parallel GNN predicts protein-ligand interactions with high accuracy.

problem Accurate prediction of protein-ligand interactions for drug design.
method Parallel Graph Neural Networks (GNN) integrating 3D structural data.
result GNN achieves high accuracy in predicting binary interactions and activity.

New method maps protein sequences to embeddings encoding structural information.

problem Inferring structural properties from amino acid sequences when structures are unknown.
method Representation learning using bidirectional LSTM models with structural similarity and residue contact maps.
result Trained embeddings improve structural similarity prediction and transfer to other tasks.

A new method predicts protein functions using variable-length sequences.

problem Computational methods for protein function prediction are slow and inaccurate for long sequences.
method Two feature sets: single fixed-sized segments and multi-sized segments, using bi-directional LSTM. Combined with MLDA features.
result Significant improvement in accuracy for long protein sequences.

Study local Weyl law on hyperbolic surfaces, identifying geodesic loops.

problem Understanding the variance of a local Weyl law on hyperbolic surfaces.
method Explicit integration of test functions, stationary phase arguments, and geometric analysis of geodesic loops.
result Identifies length-minimizing geodesic loops and sequences, proving they are simple.

Hidden semi-Markov models (HSMMs) are latent variable models which allow latent state persistence and can be viewed as a generalization of the popular hidden Markov models (HMMs). In this paper, we introduce a novel spectral algorithm to perform inference in HSMMs. Unlike expectation maximization (EM), our approach cor…

2014-07-12abs ↗pdf ↗

Graph-structured data appears frequently in domains including chemistry, natural language semantics, social networks, and knowledge bases. In this work, we study feature learning techniques for graph-structured inputs. Our starting point is previous work on Graph Neural Networks (Scarselli et al., 2009), which we modif…

2015-11-17abs ↗pdf ↗

Deep neural network translates math formula images to LaTeX sequences.

problem Translating math formula images to LaTeX sequences accurately and efficiently.
method Encoder-decoder architecture with CNN and LSTM, sequence-level training with policy gradient.
result State-of-the-art performance on sequence-based and image-based evaluation metrics.

Paper proves convergence of MDL to Einstein-Hilbert with boundary term.

problem Proving convergence of discrete MDL to continuous Einstein-Hilbert action.
method Proves \(Γ\)-convergence using diffeomorphism-natural discrete MDL-type functional.
result Identifies Carathéodory densities and obtains \(\liminf/\limsup\) bounds.

In a 1967 paper, Banchoff stated that a certain type of polyhedral curvature, that applies to all finite polyhedra, was zero at all vertices of an odd-dimensional polyhedral manifold; one then obtains an elementary proof that odd-dimensional manifolds have zero Euler characteristic. In a previous paper, the author defi…

2003-10-30abs ↗pdf ↗

The paper proposes a generic learning method for complex structures.

problem Learning algorithms for complex structures are inefficient and require significant effort.
method Mapping any complex structure onto a generic form (serialization) and applying sequence-based density estimators.
result The method is competitive or better than specialized algorithms for given structures and provides protection from overfitting.

This paper studies when particle filtering is efficient for planning in partially observed systems.

problem The efficiency of particle filtering for planning in partially observed linear dynamical systems.
method Coupling of ideal and approximate sequences to bound particle complexity.
result Polynomially many particles suffice for stable systems to approximate optimal planning.