Machine learning predicts signaling peptides from protein star graphs.
problem Predicting signaling activity of proteins from molecular structure.
method Protein star graphs, S2SNet topological indices, Machine Learning (SVM-RFE, Laplacian kernel).
result Best model predicts 98.0% signaling pathways with AUROC 0.961.
In this work, we present an application of Locally Interpretable Machine-Agnostic Explanations to 2-D chemical structures. Using this framework we are able to provide a structural interpretation for an existing black-box model for classifying biologically produced fuel compounds with regard to Research Octane Number. T…
Study improves drug prediction accuracy for pharmacokinetic parameters.
problem Limited accuracy in predicting pharmacokinetic parameters.
method Integrated transfer learning and multitask learning approach.
result Improved model generalization and predictive ability.
Transformative machine learning improves model accuracy and explainability with limited data.
problem Improving model accuracy and interpretability with limited data in scientific tasks.
method Transforming intrinsic data representations to extrinsic ones based on model predictions.
result Transformative machine learning significantly outperforms intrinsic representations in drug-design, gene expression prediction, and meta-learning.
Study compares GNNs and classical molecular featurisations for molecular property and cliff prediction.
problem Comparing GNNs and classical featurisations for molecular property and cliff prediction.
method Systematic exploration and comparison of PDVs, ECFPs, and GNNs; introduction of substructure pooling.
result Sort & Slice outperforms hash-based folding in ECFP vectorization.
Semi-supervised learning improves QSAR model predictions for novel compounds.
problem Improving model predictions for compounds not in the training set and adjusting for selection bias.
method Semi-supervised learning framework to estimate model quality and adjust for selection bias.
result Predictions for novel compounds are improved by accounting for compound similarity and selection bias.
Although artificial neural networks have occasionally been used for Quantitative Structure-Activity/Property Relationship (QSAR/QSPR) studies in the past, the literature has of late been dominated by other machine learning techniques such as random forests. However, a variety of new neural net techniques along with suc…
Multimodal deep learning improves toxicity prediction accuracy.
problem Improving prediction accuracy of chemical compound toxicity.
method Combining multiple neural network types and data representations.
result Significantly better accuracy on a toxicity benchmark.
Quantitative structure-activity relationship (QSAR) modelling is effective 'bridge' to search the reliable relationship related bioactivity to molecular structure. A QSAR classification model contains a lager number of redundant, noisy and irrelevant descriptors. To address this problem, various of methods have been pr…
QSAR models struggle to predict activity cliffs, but graph isomorphism features improve AC-sensitivity.
problem QSAR models struggle to predict activity cliffs (ACs).
method Nine distinct QSAR models combining molecular representation methods and regression techniques.
result Graph isomorphism features improve AC-sensitivity.
Method calculates financial distributions using recursive relationships.
problem Analyzing the distribution of financial functions at discrete points.
method Recursive method to calculate probability distributions.
result High accuracy demonstrated in numerical experiments.
Novel framework detects lead-lag relationships in Chinese A-share market.
problem Detecting lead-lag relationships in the Chinese A-share market.
method Two-stage framework: long-term coupling via correlation, dynamic time warping, and rank-based metrics; high-frequency data analysis via cross-correlation, Granger causality, and regression models.
result Strongly coupled stock pairs often exhibit lead-lag effects, especially at finer time scales.
Study connects manifold complexity to scalar curvature bounds.
problem Understanding the relationship between manifold complexity and scalar curvature.
method Combining quantitative operator K-theory, Lipschitz topological K-theory, and a vanishing theorem.
result Established a relationship between covering complexity and scalar curvature bounds.
TorsionNet uses reinforcement learning to efficiently generate conformers of flexible molecules.
problem Efficiently generating diverse and representative conformer sets for flexible molecules.
method Sequential conformer search technique based on reinforcement learning under the rigid rotor approximation, trained via curriculum learning.
result TorsionNet outperforms chemoinformatics methods by 4x on large branched alkanes and several orders of magnitude on biopolymer lignin.
A fundamental aspect of biological information processing is the ubiquity of sequence-function relationships -- functions that map the sequence of DNA, RNA, or protein to a biochemically relevant activity. Most sequence-function relationships in biology are quantitative, but only recently have experimental techniques f…
Study on the relationship between explanations and predictions in machine learning models.
problem Understanding the relationship between explanations and predictions in machine learning models.
method Causal inference to measure treatment effect on hyperparameters and inputs.
result The relationship between explanations and predictions is far from ideal, especially in high-performing models.
New framework models stock relationships and investor expectations for better financial market predictions.
problem Limited by predefined stock relationships and immediate effects, current financial market analysis methods need improvement.
method Jointly models investor expectations and automatically mines latent stock relationships.
result Annual return exceeds 10%, surpassing existing benchmarks.
This paper examines the quantitative finance aspects of AMMs in decentralized finance.
problem Understanding the mathematical and financial underpinnings of AMMs.
method Review of existing literature and analysis of mathematical aspects.
result Interesting relationship between AMMs and derivatives pricing and hedging.
This paper analyzes the quantitative relations between stock prices and quantities of tradable stock shares in Chinese stock markets at six time points by means of Exploratory Data Analysis (EDA) method. It is found the resulting formulae have the same structure but different parameters. This paper also uses these rela…
Study predicts social relationships using triadic influence from social networks.
problem Difficulty in quantifying social relationships and their dynamics.
method Real social networks of 13 schools, neural networks, high-dimensional embedding.
result Triadic influence achieves highest accuracy in predicting student relationships.
Single-cell gene expression data provide invaluable resources for systematic characterization of cellular hierarchy in multi-cellular organisms. However, cell lineage reconstruction is still often associated with significant uncertainty due to technological constraints. Such uncertainties have not been taken into accou…
Optimizes PnL using linear signals in quantitative finance.
problem Maximizing profit and loss in financial trading.
method Unsupervised machine learning approach that maximizes Sharpe Ratio through linear relationships and parameter optimization.
result Empirical validation and effectiveness of the model on U.S. Treasury ETF.
We study the localization of a cluster of activated vertices in a graph, from adaptively designed compressive measurements. We propose a hierarchical partitioning of the graph that groups the activated vertices into few partitions, so that a top-down sensing procedure can identify these partitions, and hence the activa…
Study finds resolution impacts human classification performance in MNIST data.
problem Understanding factors affecting human classification performance in machine learning.
method Empirical study of MNIST data at various resolutions.
result Derived a quantitative relationship between resolution and human classification performance.
Quantformer uses transformer to predict stock returns, outperforming traditional strategies.
problem Predicting stock returns in a dynamic financial market.
method Transfer learning from sentiment analysis to build investment factors using a transformer-based neural network.
result Quantformer outperforms other 100-factor-based quantitative strategies in predicting stock trends.
Paper tackles domain generalization by minimizing domain-based covariance.
problem Training data and test data have different distributions, leading to poor generalization.
method Find a central subspace minimizing domain-based covariance while preserving functional relationships.
result The proposed method achieves better generalization performance on unseen test datasets.
Why deep neural networks (DNNs) capable of overfitting often generalize well in practice is a mystery [#zhang2016understanding]. To find a potential mechanism, we focus on the study of implicit biases underlying the training process of DNNs. In this work, for both real and synthetic datasets, we empirically find that a…
The paper proves density and positive mass theorems for incomplete manifolds.
problem Proving density and positive mass theorems for manifolds with incomplete ends.
method Using harmonic asymptotics and quantitative positive mass theorem improvements.
result Improved quantitative positive mass theorem in dimensions 3 to 7.
We consider irreducible 3-manifolds M that arise as knot complements in closed 3-manifolds and that contain at most two connected strict essential surfaces. The results in the paper relate the boundary slopes of the two surfaces to their genera and numbers of boundary components. Explicit quantitative relationships, wi…
This paper finds a linear relationship between t-SNE perplexity and data set size.
problem Choosing the right perplexity for t-SNE embeddings.
method Analyzed the relationship between perplexity and data set size.
result Embeddings remain structurally consistent when perplexity is adjusted accordingly.
Error estimates found between SGD with momentum and Langevin diffusion.
problem Quantifying the difference between SGD with momentum and Langevin diffusion.
method Established error estimates using 1-Wasserstein and total variation distances.
result Quantitative error estimates between SGD with momentum and underdamped Langevin diffusion.
TDA improves understanding of B2B customer loyalty.
problem Understanding and strengthening B2B customer relationships.
method Topological Data Analysis applied to commercial data.
result TDA enhances customer base understanding and predictive model accuracy.
Study decomposes uncertainty in HK-distribution parameter estimation for QUS.
problem Uncertainty in HK-distribution parameter estimation for quantitative ultrasound.
method Bayesian Neural Networks (BNNs) for parameter estimation and uncertainty decomposition.
result Decomposes total predictive uncertainty into epistemic and aleatoric components.
Study explores factors influencing saving behavior among Dhaka employees.
problem Factors influencing saving behavior among Dhaka employees.
method Quantitative approach with cross-sectional survey design, structured questionnaire, descriptive statistics, reliability analysis, regression analysis.
result Only financial management practices had a significant positive relationship with saving behavior.
Similarity algebra extends algebraic structures with quantitative bounds.
problem Exact algebraic structures with strict axioms.
method Framework for approximate algebraic and Lie structures with ε-estimates. result Similarity structures converge to classical algebraic objects as εightarrow0. The paper analyzes relationships between various image models for restoration.
problem Unclear relationships among popular image models for restoration.
method Theoretical analysis and experimental study of image models.
result Improved denoising performance by combining multiple image models.
This study evaluates clustering algorithms on high-dimensional data.
problem Comparing clustering algorithms on high-dimensional datasets.
method Evaluation of K-means, DBSCAN, and Spectral Clustering using PCA, t-SNE, UMAP, and multiple metrics.
result UMAP preprocessing improves clustering quality across all algorithms, with Spectral Clustering excelling.
The paper examines how closed curves on surfaces intersect and how this intersection determines the curves.
problem Determining closed curves on surfaces based on their intersections.
method Constructing and studying k-equivalent curves, analyzing intersections with other curves. result Curves are determined by their intersections with all other curves, but non-simple curves require infinitely many intersections to distinguish.
We consider the problems of detection and localization of a contiguous block of weak activation in a large matrix, from a small number of noisy, possibly adaptive, compressive (linear) measurements. This is closely related to the problem of compressed sensing, where the task is to estimate a sparse vector using a small…
sPortfolio visualizes stock portfolios and factor data for better investment analysis.
problem Insufficient intuitive visual analytics for multi-factor stock portfolios.
method Develops a holistic visualization system for risk-factor, multiple-portfolios, and single-portfolios.
result Facilitates actionable insights and market trend understanding through intuitive visual analytics.
Graph neural networks improve odor prediction from molecular structure.
problem Predicting odor from molecular structure is challenging and important.
method Used graph neural networks for QSOR modeling.
result Graph neural networks significantly outperform prior methods on a novel data set.
In this paper are presented methods of impact analysis on informatics system security accidents, qualitative and quantitative methods, starting with risk and informational system security definitions. It is presented the relationship between the risks of exploiting vulnerabilities of security system, security level of …
Learning Markov blanket (MB) structures has proven useful in performing feature selection, learning Bayesian networks (BNs), and discovering causal relationships. We present a formula for efficiently determining the number of MB structures given a target variable and a set of other variables. As expected, the number of…
ATD measures language distance using neural models, recovering linguistic groupings.
problem Lack of a unified quantitative measure for cross-linguistic distance.
method Pretrained multilingual language models, attention mechanisms, optimal transport.
result ATD quantifies representational distance between languages, recovering linguistic groupings.
Investment strategy using fractional Kelly portfolios for better growth expectations.
problem Understanding optimal growth strategies for investors with varying risk appetites.
method Developed a mathematical framework for fractional-Kelly portfolios, analyzing Sharpe ratios and log-returns.
result Fractional Kelly portfolios provide a simple distributional relationship between Sharpe ratio, fractional coefficient, and log-returns.
SpArX creates faithful explanations of neural networks' decision-making.
problem Challenges in explaining neural networks' decisions.
method Sparsifies MLPs while maintaining structure, then translates into QAFs for argumentative explanations.
result SpArX provides more faithful explanations than existing methods.
The study analyzes stock price dynamics of S&P 100 using 250k data points.
problem Characterize investor behavior and price dynamics in S&P 100 stocks.
method Two-way fixed-effects model applied to 250k data points of S&P 100 stocks.
result Uncovered nonlinear relationship between return and trend, quantifying trader motivations.
We present a new approach to understanding credit relationships between commercial banks and quoted firms, and with this approach, examine the temporal change in the structure of the Japanese credit network from 1980 to 2005. At each year, the credit network is regarded as a weighted bipartite graph where edges corresp…