Before retiring, looking back to forty years of writing and publishing scientific papers, I decided to present to the scientific community a selection of my scientific works. I chose mostly articles published in prestigious journals or Proceedings that made a certain impact in the scientific world. I have selected thir…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper explores MMPR to select diverse models for scientific insight.
iKF method uncovers complex variable interactions for scientific discovery.
New method speeds up model selection for complex scientific tasks.
Framework evaluates AI proposals for drug discovery, finds no LLM advantage.
Practical or scientific considerations often lead to selecting a subset of parameters as ``important.'' Inferences about those parameters often are based on the same data used to select them in the first place. That can make the reported uncertainties deceptively optimistic: confidence intervals that ignore selection g…
In the era of big data, analysts usually explore various statistical models or machine learning methods for observed data in order to facilitate scientific discoveries or gain predictive power. Whatever data and fitting procedures are employed, a crucial step is to select the most appropriate model or method from a set…
Scientific Computing relies on executing computer algorithms coded in some programming languages. Given a particular available hardware, algorithms speed is a crucial factor. There are many scientific computing environments used to code such algorithms. Matlab is one of the most tremendously successful and widespread s…
CausalGame benchmarks LLM agents' causal thinking in games.
Survey of deep learning models for scientific discovery.
Develops a Bayesian framework for symbolic regression of scientific expressions.
A new method selects features for ERGMs to improve network modeling.
In recent years, ideas from statistics and scientific computing have begun to interact in increasingly sophisticated and fruitful ways with ideas from computer science and the theory of algorithms to aid in the development of improved worst-case algorithms that are useful for large-scale scientific and Internet data an…
The increasing size and complexity of scientific data could dramatically enhance discovery and prediction for basic scientific applications. Realizing this potential, however, requires novel statistical analysis methods that are both interpretable and predictive. We introduce Union of Intersections (UoI), a flexible, m…
The process of collecting and organizing sets of observations represents a common theme throughout the history of science. However, despite the ubiquity of scientists measuring, recording, and analyzing the dynamics of different processes, an extensive organization of scientific time-series data and analysis methods ha…
GSR optimizes tasks in scientific workflows, improving performance across diverse applications.
PRISM infers model structures and parameters from simulations, controlling complexity at test time.
This paper provides a guide to feature importance methods for better scientific inference.
FreB protocol uses AI to infer hidden parameters with valid confidence regions.
LLMs fail to match statistical ground truth despite stable run-to-run performance.
Feature selection is central to contemporary high-dimensional data analysis. Grouping structure among features arises naturally in various scientific problems. Many methods have been proposed to incorporate the grouping structure information into feature selection. However, these methods are normally restricted to a li…
Machine learning algorithms such as linear regression, SVM and neural network have played an increasingly important role in the process of scientific discovery. However, none of them is both interpretable and accurate on nonlinear datasets. Here we present contextual regression, a method that joins these two desirable …
Interpretable classification models are built with the purpose of providing a comprehensible description of the decision logic to an external oversight agent. When considered in isolation, a decision tree, a set of classification rules, or a linear model, are widely recognized as human-interpretable. However, such mode…
Refining one's hypotheses in the light of data is a common scientific practice; however, the dependency on the data introduces selection bias and can lead to specious statistical analysis. An approach for addressing this is via conditioning on the selection procedure to account for how we have used the data to generate…
In this paper we review the concepts of Bayesian evidence and Bayes factors, also known as log odds ratios, and their application to model selection. The theory is presented along with a discussion of analytic, approximate and numerical techniques. Specific attention is paid to the Laplace approximation, variational Ba…
Bayesian framework integrates spectral deconvolution with expert reasoning for robust peak estimation.
Polynomial chaos surrogates quantify epistemic uncertainty in AI-driven scientific models.
NGMs create mirrored features to assess neural network feature importance.
The paper develops a test for independence of selected Gaussian variables after thresholding correlations.
Stability is an important aspect of a classification procedure because unstable predictions can potentially reduce users' trust in a classification system and also harm the reproducibility of scientific conclusions. The major goal of our work is to introduce a novel concept of classification instability, i.e., decision…
Two-sample feature selection is the problem of finding features that describe a difference between two probability distributions, which is a ubiquitous problem in both scientific and engineering studies. However, existing methods have limited applicability because of their restrictive assumptions on data distributoins …
Machine learning improves network classification and model selection.
Gryffin optimizes categorical variables in materials design, leveraging expert knowledge.
In neuroimaging data analysis, Gaussian graphical models are often used to model statistical dependencies across spatially remote brain regions known as functional connectivity. Typically, data is collected across a cohort of subjects and the scientific objectives consist of estimating population and subject-specific g…
Online method selects candidates from data streams, ensuring irreversible decisions.
Proposes a method for stable variable selection in high-dimensional data.
A method selects candidates based on predictions with statistical control.
New architectures improve KANs, making them more interpretable and accurate.
Understanding how features interact with each other is of paramount importance in many scientific discoveries and contemporary applications. Yet interaction identification becomes challenging even for a moderate number of covariates. In this paper, we suggest an efficient and flexible procedure, called the interaction …
Machine learning finds new natural laws from noisy data.
New method controls false edge detections in Gaussian graphical models.
A new method selects features for clustering without labels.
FAStEN efficiently selects features in high-dimensional functional data.
New methods for selecting variables in complex biomedical data.
Simultaneous inference after model selection is of critical importance to address scientific hypotheses involving a set of parameters. In this paper, we consider high-dimensional linear regression model in which a regularization procedure such as LASSO is applied to yield a sparse model. To establish a simultaneous pos…
By mapping the most advanced elements of the contemporary social interactions, the world scientific collaboration network develops an extremely involved and heterogeneous organization. Selected characteristics of this heterogeneity are studied here and identified by focusing on the scientific collaboration community of…
Method selects interpretable circular coordinates from data.
xVal tokenizes numbers continuously for better scientific model training.