A new pruning method reduces DNN size and interconnectivity using brain network principles.
problem Over-parameterization in deep neural networks causes memory and hardware cost issues.
method Structural pruning scheme based on Small-World model, trimming network before training.
result Reduced model size by 2.3% on LeNet-5 for MNIST and 9.02% on VGG-16 for CIFAR-10.
SWNets optimize DL architectures for faster convergence.
problem Excessive training parameters in deep learning models.
method Transforms network topology to reach Small-World Network boundary.
result SWNets achieve faster convergence with fewer parameters.
Improved tabular models learn better from real-world data.
problem Tabular models perform poorly on real-world datasets when trained only on synthetic data.
method Continued pre-training on a curated set of real-world datasets.
result Real-TabPFN achieves superior predictive accuracy on 29 datasets.
Unified algorithm for latent patterns in SBM and SWM models.
problem Invalid analysis due to misspecified models in graph analysis.
method Combining kernel learning, spectral graph theory, and dimensionality reduction.
result First statistically sound polynomial-time algorithm for latent patterns.
Estimates CATEs for structured treatments using a new decomposition method.
problem Estimating conditional average treatment effects for complex data types.
method Generalized Robinson decomposition, isolating causal estimand, arbitrary model plugging, quasi-oracle convergence guarantee.
result Demonstrates superior performance in CATE estimation compared to prior work.
SHADOWCAST generates graphs with user-specified attributes.
problem Controlling graph generation with understandable structures.
method Conditional generative adversarial network guided by Markov model.
result Competitive performance in generating desired graphs.
RL controls small soccer robots in a real league, beating human-designed policies.
problem Training robots to play complex, real-world sports.
method Sim-to-Real RL approach, training in simulated environment, applying to real-world robots.
result Robots learned policies to compete effectively, beating human-designed strategies.
Novel approach for SEM in small samples with p>n.
problem Small sample size and p>n issues in factor-based SEM. method Reformulates covariance structure into self-covariance and cross-covariance, defines a feasible set with relative error constraint.
result Improved stability and directional information in small-sample settings.
Framework explains deep learning generalization by comparing real and ideal worlds.
problem Understanding why deep models generalize well in practice.
method Integrates real-world empirical loss with ideal population loss to decompose test error.
result The gap between real and ideal worlds is small in deep learning, suggesting robust optimization leads to good generalization.
Augment small datasets with synthetic backgrounds to train lightweight CNNs for human pose estimation.
problem Training CNNs from limited real-world data for human pose estimation.
method Synthetic background substitution for data augmentation.
result Improves generalization to unseen environments.
Study on limits of community detection in various network models.
problem Limits of community detection in network models.
method Analysis of several network models including Stochastic Block Model, Exponential Random Graph Model, Latent Space Model, Directed Preferential Attachment Model, and Directed Small-world Model.
result Information-theoretic limits for recovery of node labels in network models.
The problem of content search through comparisons has recently received considerable attention. In short, a user searching for a target object navigates through a database in the following manner: the user is asked to select the object most similar to her target from a small list of objects. A new object list is then p…
Dreamer 4 learns Minecraft tasks from videos alone.
problem Accurately predicting object interactions in complex environments.
method Reinforcement learning inside a fast, accurate world model.
result Dreamer 4 outperforms previous models in Minecraft, learning from only offline data.
The study examines price formation in complex networks and finds efficiency varies by network structure.
problem Understanding price formation and efficiency in complex networks.
method Price formation experiments with human subjects in large networks, agent-based model construction.
result Prices are higher and trade less efficient in small-world networks compared to random networks.
The study examines how prior and likelihood choices affect Bayesian matrix factorisation on small datasets.
problem Improving predictive performance of Bayesian matrix factorisation on small datasets.
method Review and comparison of 16 Bayesian matrix factorisation models across four groups: Gaussian-likelihood with real-valued priors, nonnegative priors, semi-nonnegative models, and Poisson-likelihood approaches.
result Poisson models give poor predictions, and nonnegative models are more constrained than real-valued ones.
Study analyzes Bitcoin and Bitcoin Cash networks for small-world properties.
problem Understanding the structure and behavior of Bitcoin and Bitcoin Cash networks.
method Network analysis of Bitcoin and Bitcoin Cash networks to evaluate small-world properties.
result Found that networks exhibit small-world behavior, suggesting mechanisms leading to current structure.
Modern computer vision algorithms typically require expensive data acquisition and accurate manual labeling. In this work, we instead leverage the recent progress in computer graphics to generate fully labeled, dynamic, and photo-realistic proxy virtual worlds. We propose an efficient real-to-virtual world cloning meth…
Technique creates highly accurate small models for better interpretability.
problem Balancing model accuracy and interpretability for constrained models.
method Identifies optimal training distribution for a given model size using Infinite Mixture Model with Beta components and Bayesian Optimization.
result Significant improvements in F1-score, up to 100% in some cases.
This study applies EMD to MSCI World index and converts IMFs into graphs for GNN modeling.
problem Modeling financial time series with GNNs.
method EMD, CEEMDAN, graph transformations (natural visibility, horizontal visibility, recurrence, transition graphs), topological analysis.
result High-frequency IMFs yield dense, highly connected small-world graphs; low-frequency IMFs produce sparser networks.
Empirical data of supermarket sales show stylised facts that are similar to stock markets, with a broad (truncated) Levy distribution of weekly sales differences in the baseline sales [R.D. Groot, Physica A 353 (2005) 501]. To investigate the cause of this, the influence of social interactions and advertisements are st…
NPGNN improves graph link prediction by adapting to new graphs.
problem Inductive link prediction in graphs with limited training data.
method Meta-learning with graph neural networks (NPGNN).
result NPGNN outperforms state-of-the-art models in real-world graphs.
Improved Naive Bayes for text classification with small datasets.
problem Poor performance of Naive Bayes in small training datasets.
method Introducing a correlation factor to Naive Bayes estimator.
result Our method achieves better accuracy than traditional Naive Bayes.
This comment reexamines Simard et al.'s work in [D. Simard, L. Nadeau, H. Kroger, Phys. Lett. A 336 (2005) 8-15]. We found that Simard et al. calculated mistakenly the local connectivity lengths Dlocal of networks. The right results of Dlocal are presented and the supervised learning performance of feedforward neural n…
Empirical model tackles decision problems without specifying states of the world.
problem Decision problems under uncertainty with inaccessible states of the world.
method Empirical approach using observed act--consequence pairs as model primitives.
result Optimality in empirical decision problems addressed using protocol-based empirical choice functions.
Framework for applying GPs to real-world data with scalability guidelines.
problem Deployment of Gaussian Processes (GPs) is hindered by computational costs and lack of guidelines.
method Proposed a framework for identifying GP suitability and setting up robust models, formalizing decisions of experienced practitioners.
result More accurate results at test time for glacier elevation change case study.
We provide approximations for VIX futures and options in forward variance models.
problem Modeling VIX futures and options in forward variance models.
method Weak approximations and explicit formula derivation for VIX futures and options.
result Explicit combinations of Black-Scholes prices and greeks for option price approximations.
Modeling reinsurance network contagion and its risks.
problem Contagion risk in reinsurance networks underestimates simpler models.
method Developed a model for reinsurance network contagion, characterized fixed points, and developed algorithms for computation.
result Reinsurance networks are highly sensitive to parameters and network structure, leading to significant losses.
Datasets containing large samples of time-to-event data arising from several small heterogeneous groups are commonly encountered in statistics. This presents problems as they cannot be pooled directly due to their heterogeneity or analyzed individually because of their small sample size. Bayesian nonparametric modellin…
FlowMO uses Gaussian Processes for molecular property prediction with uncertainty.
problem Predicting molecular properties with uncertainty for small datasets.
method Gaussian Processes implemented in FlowMO, built on GPflow and RDKit.
result Comparable predictive performance to deep learning but superior uncertainty calibration.
MILABOT is a chatbot trained to converse with humans using deep reinforcement learning.
problem Developing a conversational agent capable of handling open-domain small talk topics.
method Deep reinforcement learning applied to crowdsourced and real-world data.
result MILABOT outperformed other systems in A/B testing with real-world users.
Paper explains why small-loss criterion works for learning from noisy labels.
problem Learning from noisy labels in deep learning with limited labeled data.
method Theoretical analysis and reformulation of the small-loss criterion.
result Theoretical explanation and reformulation of the small-loss criterion.
Proposes a new signal model for high-dimensional, small-sample-size data.
problem Signal detection in high-dimensional, small-sample-size datasets.
method Intrinsic signal model based on dynamical system assumption.
result Taguchi method effectively detects signals in the proposed model.
The equity risk premium is derived from SPX option chains using a model-light approach.
problem Estimating the equity risk premium from option data.
method Model-light approach using Gaussian mixture models and exponential tilting.
result The equity risk premium is calculated from the real-world probability densities inferred from option quotes.
Variational Bayesian inference and (collapsed) Gibbs sampling are the two important classes of inference algorithms for Bayesian networks. Both have their advantages and disadvantages: collapsed Gibbs sampling is unbiased but is also inefficient for large count values and requires averaging over many samples to reduce …
Improves probability estimates for small datasets in multi-class problems.
problem Inaccurate probability estimates in classification tasks, especially on small datasets.
method Introduced Data Generation and Grouping algorithm to improve calibration on small datasets, then applied to multi-class problems.
result Calibration error can be decreased using the proposed approach.
Machine learning improves network classification and model selection.
problem Quantifying suitability of generative models for network structures.
method Interpretable machine learning to classify simulated networks based on features and interactions.
result Specific network features and their interactions are crucial for distinguishing generative models.
Study uses Apple ML to accurately detect and classify lung cancer.
problem Accurate diagnosis and sub-classification of non-small cell lung cancer.
method Evaluation of Apple Create ML module on histopathological images.
result 100% detection and successful subclassification of non-small cell lung cancer.
Risk scores are simple classification models that let users make quick risk predictions by adding and subtracting a few small numbers. These models are widely used in medicine and criminal justice, but are difficult to learn from data because they need to be calibrated, sparse, use small integer coefficients, and obey …
Machines, not humans, are the world's dominant knowledge accumulators but humans remain the dominant decision makers. Interpreting and disseminating the knowledge accumulated by machines requires expertise, time, and is prone to failure. The problem of how best to convey accumulated knowledge from computers to humans i…
SFM resolves small-scale physics challenges in weather data.
problem Challenges in super-resolving small-scale details in physical sciences like weather.
method Encoding inputs to a latent base distribution, flow matching for stochastic details, adaptive noise scaling.
result SFM framework significantly outperforms existing methods.
Simple private estimators for mean and covariance outperform existing methods.
problem Private estimation of mean and covariance at small sample sizes.
method Differentially private estimators for multivariate sub-Gaussian data.
result Asymptotic error rates match theoretical bounds and outperform previous methods.
Improves natural accuracy of deep learning models by combining robust predictions and features.
problem Maintaining natural accuracy while resisting adversarial attacks.
method Ensemble methods combining robust and standard models.
result Optimized natural accuracy through ensemble of robust models.
RoPE framework calibrates misspecified simulators for reliable inference.
problem Misspecification compromises reliability of simulation-based inference.
method Data-driven calibration using optimal transport and a small calibration set.
result RoPE framework improves inference accuracy and uncertainty calibration.
Differential calculus on metric spaces is contained in the algebraic study of normed groupoids with δ-structures. Algebraic study of normed groups endowed with dilatation structures is contained in the differential calculus on metric spaces. Thus all algebraic properties of the small world of normed groups with dilat…
TAnoGan detects anomalies in time series data using GANs.
problem Anomaly detection in time series data.
method Generative Adversarial Networks (GAN) for unsupervised anomaly detection.
result TAnoGan outperforms traditional and neural network models in anomaly detection.
FedFaiREE addresses fairness in decentralized learning with small samples.
problem Ensuring fairness in decentralized federated learning with limited data.
method FedFaiREE is a post-processing algorithm for distribution-free fair learning in decentralized settings with small samples.
result FedFaiREE provides theoretical guarantees for both fairness and accuracy in decentralized environments.
We perform a stability analysis for the utility maximization problem in a general semimartingale model where both liquid and illiquid assets (random endowments) are present. Small misspecifications of preferences (as modeled via expected utility), as well as views of the world or the market model (as modeled via subjec…
Study shows disentanglement models learn correlations from data, impacting fairness.
problem Disentanglement models learn correlations in real-world data, affecting downstream applications.
method Empirical study on 4260 models, analyzing correlations in latent representations.
result Systematically induced correlations are learned by disentanglement models, impacting fairness.