Simpler model outperforms deep learning for disease prediction.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Tree ensembles, such as random forest and boosted trees, are renowned for their high prediction performance, whereas their interpretability is critically limited. In this paper, we propose a post processing method that improves the model interpretability of tree ensembles. After learning a complex tree ensembles in a s…
Simpler linear models outperform complex GCN encoders for graph tasks.
New criteria distinguish cause from effect in data, overcoming statistical limitations.
We introduce a method, KL-LIME, for explaining predictions of Bayesian predictive models by projecting the information in the predictive distribution locally to a simpler, interpretable explanation model. The proposed approach combines the recent Local Interpretable Model-agnostic Explanations (LIME) method with ideas …
A hybrid model combines piecewise linear and neural components for interpretable predictions.
Noise increases the Rashomon ratio, leading simpler models to perform similarly to complex ones.
A new method creates simpler, more interpretable decision trees from complex ensembles.
Improved model for multivariate time series prediction with simpler architecture.
GRUwE improves irregular time series prediction with simpler, efficient RNN-based approach.
Bayes posterior yields worse predictions than simpler methods in deep neural networks.
New classifiers converge under large data, simplifying complex models.
The lack of interpretability remains a key barrier to the adoption of deep models in many applications. In this work, we explicitly regularize deep models so human users might step through the process behind their predictions in little time. Specifically, we train deep time-series models so their class-probability pred…
This paper re-examines conformal e-prediction and its advantages over conformal prediction.
BusTr predicts bus travel times from real-time traffic forecasts.
A simpler proof for apex graphs in McCarty and Thomas' conjecture.
State of the art machine learning algorithms are highly optimized to provide the optimal prediction possible, naturally resulting in complex models. While these models often outperform simpler more interpretable models by order of magnitudes, in terms of understanding the way the model functions, we are often facing a …
The paper combines Bitcoin price models with expert corrections for better predictions.
Material scientists are increasingly adopting the use of machine learning (ML) for making potentially important decisions, such as, discovery, development, optimization, synthesis and characterization of materials. However, despite ML's impressive performance in commercial applications, several unique challenges exist …
Making an adaptive prediction based on one's input is an important ability for general artificial intelligence. In this work, we step forward in this direction and propose a semi-parametric method, Meta-Neighborhoods, where predictions are made adaptively to the neighborhood of the input. We show that Meta-Neighborhood…
Simpler proof for Kielak's virtual fibering criterion.
Simpler algorithm learns shallow networks faster.
We construct flexible likelihoods for multi-output Gaussian process models that leverage neural networks as components. We make use of sparse variational inference methods to enable scalable approximate inference for the resulting class of models. An attractive feature of these models is that they can admit analytic pr…
Investor sentiment improves model accuracy but complexity doesn't always boost predictive power.
METRO predicts reactions using minimal templates, reducing computational overhead and achieving state-of-the-art results.
In statistical relational learning, the link prediction problem is key to automatically understand the structure of large knowledge bases. As in previous studies, we propose to solve this problem through latent factorization. However, here we make use of complex valued embeddings. The composition of complex embeddings …
New model predicts video sequences with latent dynamics.
Neural Architecture Search methods are effective but often use complex algorithms to come up with the best architecture. We propose an approach with three basic steps that is conceptually much simpler. First we train N random architectures to generate N (architecture, validation accuracy) pairs and use them to train a …
Better stock market predictions can be made by simpler topic models.
New approach for algorithms that learn predictors to improve performance.
We study the performance of Local Causal Discovery (LCD), a simple and efficient constraint-based method for causal discovery, in predicting causal effects in large-scale gene expression data. We construct practical estimators specific to the high-dimensional regime. Inspired by the ICP algorithm, we use an optional pr…
Researchers use estimated Kolmogorov complexity for better link prediction in graphs.
CPPS selects portfolios using conformal prediction for better returns.
Study compares deep learning models for volatility prediction using multivariate data.
Networks are powerful data structures, but are challenging to work with for conventional machine learning methods. Network Embedding (NE) methods attempt to resolve this by learning vector representations for the nodes, for subsequent use in downstream machine learning tasks. Link Prediction (LP) is one such downstream…
Deep learning models, especially CNNs, can predict radio frequency power faster than traditional methods.
Simplifies neural regression by combining two sub-networks for predictions and uncertainties.
We define a new Hurwitz problem which is essentially a small core of the simple Hurwitz problem. The corresponding Hurwitz numbers have simpler formulae, satisfy effective recursion relations and determine the simple Hurwitz numbers. We also apply this idea of finding a smaller simpler enumerative problem to orbifold H…
We provide an alternative, simpler proof of the existence of thick triangulations for noncompact manifolds. Moreover, this proof is simpler than the original one given in \cite{pe}, since it mainly uses tools of elementary differential topology. The role played by curvatures in this construction is also…
We present a new family of exchangeable stochastic processes, the Functional Neural Processes (FNPs). FNPs model distributions over functions by learning a graph of dependencies on top of latent representations of the points in the given dataset. In doing so, they define a Bayesian model without explicitly positing a p…
When faced with the problem of learning a model of a high-dimensional environment, a common approach is to limit the model to make only a restricted set of predictions, thereby simplifying the learning problem. These partial models may be directly useful for making decisions or may be combined together to form a more c…
BYOL-Explore learns to explore visually-rich environments by predicting world dynamics.
We review statistical properties of models generated by the application of a (positive and negative order) fractional derivative operator to a standard random walk and show that the resulting stochastic walks display slowly-decaying autocorrelation functions. The relation between these correlated walks and the well-kno…
Reduces constructing multiplicative connections to simpler tasks.
We show how to control the generalization error of time series models wherein past values of the outcome are used to predict future values. The results are based on a generalization of standard i.i.d. concentration inequalities to dependent data without the mixing assumptions common in the time series setting. Our proo…
Stochastic gradient methods are dominant in nonconvex optimization especially for deep models but have low asymptotical convergence due to the fixed smoothness. To address this problem, we propose a simple yet effective method for improving stochastic gradient methods named predictive local smoothness (PLS). First, we …
SGD tends to favor simpler subnetworks, improving generalization.
The transport literature is dense regarding short-term traffic predictions, up to the scale of 1 hour, yet less dense for long-term traffic predictions. The transport literature is also sparse when it comes to city-scale traffic predictions, mainly because of low data availability. In this work, we report an effort to …