Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

0.8%1.6%2.4%3.2% · Feb 199719922001200920172026
48 results for plant transcriptomics

PLIT identifies plant lncRNAs from RNA-seq data with high accuracy.

problem Inaccurate identification of lncRNAs in plant transcriptomic datasets.
method PLIT uses L1 regularization and iRF classification to select optimal features from sequence and codon-bias data.
result PLIT outperforms existing CPC tools in identifying lncRNAs in plant RNA-seq datasets.

A model to fill in missing gene data from spatial studies and scRNA-seq.

problem Imputing missing gene expression measurements from spatial transcriptomics.
method A deep generative model (gimVI) for integrating spatial transcriptomic and scRNA-seq data.
result gimVI outperforms existing methods in imputing missing genes.

This study benchmarks transcriptomics models for perturbation analysis, finding scVI and PCA superior.

problem Limited evaluation of transcriptomics foundation models for perturbation analysis.
method Developed a novel evaluation framework using diverse public datasets from different sequencing techniques and cell lines.
result scVI and PCA identified as superior models for understanding biological perturbations.

Sparse neural networks visualize paired transcriptomic and electrophysiological data.

problem Efficiently analyzing and visualizing paired multivariate neuroscientific data.
method Sparse deep neural networks with a two-dimensional bottleneck and group lasso penalty.
result Biologically interpretable two-dimensional visualizations of paired data.

Generative model tailors anticancer drugs based on transcriptomic data.

problem Designing effective anticancer drugs considering genetic profiles.
method RL framework using pretrained VAEs to generate compounds conditioned on transcriptomic data.
result Generative model produces molecules with high predicted inhibitory effects.

TransST improves spatial transcriptomics data analysis by identifying cell clusters and biomarkers.

problem Low resolution and insufficient sequencing depth in spatial transcriptomics data.
method Transfer learning framework to adaptively leverage external cell-labeled information.
result TransST successfully identifies five biologically meaningful cell clusters and separates adipose tissues from connective issues.

NO-BEARS algorithm speeds up gene network inference from transcriptomic data.

problem Constructing accurate gene regulatory networks from transcriptomic data.
method NO-BEARS algorithm, based on NOTEARS, with new constraint and polynomial regression loss.
result Significantly reduced computational time and improved accuracy in inferring gene regulatory networks.

A new parallel clustering method improves speed and accuracy for single cell transcriptomic data.

problem Challenges in clustering single cell transcriptomic data, including poor quality, lack of prior knowledge, and slow computation.
method Parallel Split Merge Sampling on Dirichlet Process Mixture Model (Para-DPMM).
result The Para-DPMM model outperforms existing methods in clustering quality and computational speed.

Exact partitioning of high-order planted models achieved through convex optimization.

problem Efficiently partitioning hypergraphs generated by high-order planted models.
method Solving a computationally efficient convex optimization problem with a tensor nuclear norm constraint.
result Exact recovery of true underlying cluster structures with high probability.

Study of SK-N-AS cells' response to methamidophos using transcriptomics.

problem Understanding the transcriptional response of SK-N-AS cells to methamidophos exposure.
method Combination of statistical and machine learning methods for anomaly detection and causal network inference.
result Identification of key processes and transcripts involved in the response to methamidophos.

AI system synthesizes chemical plant operation procedures for efficiency and stability.

problem Developing efficient and stable operation procedures for complex chemical plants.
method Integrates automated reasoning, deep reinforcement learning, and dynamic simulation with external knowledge.
result Synthesized procedure achieves faster recovery from malfunctions compared to standard PID control.

Tackles the computational hardness of HPC detection, conjecturing equivalence to PC detection.

problem Computational hardness of hypergraphic planted clique detection.
method No specific method mentioned; focuses on conjecturing equivalence.
result Equivalence of computational hardness between HPC and PC detection.

Deep learning identifies transcriptomic patterns and cell types associated with SARS-CoV-2 infection and COVID-19 severity.

problem Understanding how SARS-CoV-2 varies in infecting and causing severe COVID-19.
method Developed a new approach to generating self-supervised edge features, using Graph Attention Networks (GAT) and Set Transformer.
result Achieved state-of-the-art performance in predicting disease state of individual cells using single-cell RNA sequencing data.

Plants monitor their surrounding environment and control their physiological functions by producing an electrical response. We recorded electrical signals from different plants by exposing them to Sodium Chloride (NaCl), Ozone (O3) and Sulfuric Acid (H2SO4) under laboratory conditions. After applying pre-processing tec…

2017-05-13abs ↗pdf ↗

Study connects database alignment and planted matching using Gaussian features.

problem Identify matching between correlated user features in anonymized databases.
method Derived results for database alignment and planted matching, showing connections and thresholds.
result Performance thresholds for database alignment converge to planted matching when feature dimensionality is sufficiently high.

A new method infers causal gene regulatory networks from parallel CRISPR interventions and transcriptomic data.

problem Learning causal gene regulatory networks from observational data is complicated by lack of identifiability and a combinatorial solution space.
method A continuous optimization framework that leverages observational and interventional data to infer a single causal structure, assuming a linear Structural Equation Model (SEM).
result A provably consistent estimator of the true DAG under mild assumptions.

Random Planted Forest interprets tree-based models by keeping some splits, leading to more interpretable predictions.

problem Estimating the unknown regression function from lower-order interaction terms.
method Modifying the random forest algorithm by keeping certain leaves instead of deleting them, resulting in non-binary trees called planted trees.
result The random planted forest achieves asymptotically optimal convergence rates up to a logarithmic factor when the interaction bound is low.

Study information limits for community detection in sub-hypergraphs.

problem Identify limits for exact community detection in sub-hypergraphs.
method Use Fano's inequality to define model parameters and identify success and failure regions.
result Identify regions where algorithms succeed or fail in exact recovery.

We consider the problem of online learning of optimal control for repeatedly operated systems in the presence of parametric uncertainty. During each round of operation, environment selects system parameters according to a fixed but unknown probability distribution. These parameters govern the dynamics of a plant. An ag…

2016-09-18abs ↗pdf ↗

Paper shows statistical-computational gaps in learning sparse mixtures and robust estimation.

problem Statistical-computational gaps in learning sparse mixtures and robust estimation.
method Average-case reduction techniques, Imbalanced Sparse Gaussian Mixtures, and algorithmic change of measure.
result New hardness results for robust sparse mean estimation, semirandom planted dense subgraph, and universality principle for sparse mixture problems.

Graph clustering involves the task of dividing nodes into clusters, so that the edge density is higher within clusters as opposed to across clusters. A natural, classic and popular statistical setting for evaluating solutions to this problem is the stochastic block model, also referred to as the planted partition model…

2012-10-11abs ↗pdf ↗

Power plant is a complex and nonstationary system for which the traditional machine learning modeling approaches fall short of expectations. The ensemble-based online learning methods provide an effective way to continuously learn from the dynamic environment and autonomously update models to respond to environmental c…

2017-10-19abs ↗pdf ↗

With the wealth of high-throughput sequencing data generated by recent large-scale consortia, predictive gene expression modelling has become an important tool for integrative analysis of transcriptomic and epigenetic data. However, sequencing data-sets are characteristically large, and previously modelling frameworks …

2015-07-21abs ↗pdf ↗

In this work we propose a method to compute continuous embeddings for kmers from raw RNA-seq data, without the need for alignment to a reference genome. The approach uses an RNN to transform kmers of the RNA-seq reads into a 2 dimensional representation that is used to predict abundance of each kmer. We report that our…

2018-10-08abs ↗pdf ↗

Paper proposes a predictive maintenance system for solar plants using big data.

problem Fault prediction in photovoltaic plants to reduce downtime and maintenance costs.
method Data-driven approach with unsupervised clustering and Pattern Recognition Neural Network.
result Effective prediction of both generic and specific faults, up to 7 days in advance.

Paper explores limits of high-order clustering with planted structures.

problem Statistical and computational limits of high-order clustering with planted structures.
method Developed methods for detection and recovery of clusters, identified signal-to-noise ratio boundaries.
result Sharp boundaries of signal-to-noise ratio for statistical and computational feasibility.

Develops a new Gaussian process method for efficient Bayesian inference of plant root parameters in the Richards equation.

problem Estimating unknown parameters in nonlinear PDEs for agricultural studies.
method Gaussian process collocation with importance sampling and Bayesian optimization.
result Our method yields robust estimates with uncertainty quantification for plant root parameters.

Interpretable neural network for plant traits and species identification.

problem Plant phenotyping and identification.
method Neural network trained on UPWINS spectral library, with visualization of weights for trait-based spectral features.
result 90% accuracy in species identification with interpretable neural network.

A neural collaborative filtering method predicts corn hybrid yield performance.

problem Predicting yield performance of untested hybrid combinations in plant breeding.
method Ensemble of matrix factorization and neural networks.
result The model significantly outperformed other models in the Syngenta Crop Challenge.

This study reviews and evaluates clustering methods for single-cell RNA-seq data.

problem Identifying and characterizing novel cell types from single-cell RNA-seq data.
method Review and performance comparison of clustering methods.
result Performance comparison experiments on two datasets.

Single-cell RNA sequencing (scRNA-seq) is a fast growing approach to measure the genome-wide transcriptome of many individual cells in parallel, but results in noisy data with many dropout events. Existing methods to learn molecular signatures from bulk transcriptomic data may therefore not be adapted to scRNA-seq data…

2018-02-26abs ↗pdf ↗

Deep learning predicts plant growth and yield in greenhouses.

problem Predicting plant growth and yield for better greenhouse management.
method Utilized a new deep recurrent neural network (RNN) with LSTM neurons to model growth parameters.
result Deep learning models outperformed traditional ML methods in predicting plant growth and yield.

The paper uses a graph autoencoder to learn unbiased plant-pollinator interaction embeddings.

problem Sampling bias in citizen science data affects ecological network analysis.
method Bipartite graph variational autoencoder with HSIC for fairness.
result The method mitigates sampling bias and provides unbiased embeddings.

Study shows it's impossible to count communities without finding them.

problem Determining the number and sizes of communities in random graph models.
method Hypothesis testing between models with different community structures, using low-degree polynomial framework.
result Testing between two different planted distributions is as hard as finding the communities.

PerturBench benchmarks ML models for cellular perturbation analysis.

problem Standardizing benchmarking in modeling single cell transcriptomic responses to perturbations.
method Modular platform, diverse datasets, metrics, extensive evaluation, rank metrics.
result Simpler models are competitive and scale well with larger datasets.