PAC-Bayesian matrix completion with a spectral scaled Student prior offers efficient inference.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Self-training with noisy student-teacher boosts keyword spotting accuracy.
Improved image reconstruction using VAEs with Student's t-prior.
Knowledge distillation is an effective technique that transfers knowledge from a large teacher model to a shallow student. However, just like massive classification, large scale knowledge distillation also imposes heavy computational costs on training models of deep neural networks, as the softmax activations at the la…
Study improves fractional posterior for 1-bit matrix completion.
Estimates model performance from compute budget for distillation.
Student- processes have recently been proposed as an appealing alternative non-parameteric function prior. They feature enhanced flexibility and predictive variance. In this work the use of Student- processes are explored for multi-objective Bayesian optimization. In particular, an analytical expression for the h…
A new method identifies a stable subnetwork in overparameterized student models.
New method improves generative modeling on convex domains using regularized mirror maps and Student-t priors.
We propose a generalized double Pareto prior for Bayesian shrinkage estimation and inferences in linear models. The prior can be obtained via a scale mixture of Laplace or normal distributions, forming a bridge between the Laplace and Normal-Jeffreys' priors. While it has a spike at zero like the Laplace density, it al…
Proposes scale mixture of NNGPs for more flexible stochastic processes.
Bayesian neural networks approximate Student-t processes in the infinite-width limit.
A VB method for high-dimensional regression with student-t priors achieves nearly optimal performance and computational efficiency.
We investigate the Student-t process as an alternative to the Gaussian process as a nonparametric prior over functions. We derive closed form expressions for the marginal likelihood and predictive distribution of a Student-t process, by integrating away an inverse Wishart process prior over the covariance kernel of a G…
Study shows human advisors use context to improve student outcomes in algorithm-assisted advising.
Tailoring the presentation of information to the needs of individual students leads to massive gains in student outcomes~\cite{bloom19842}. This finding is likely due to the fact that different students learn differently, perhaps as a result of variation in ability, interest or other factors~\cite{schiefele1992interest…
Proposes new models to predict student grades more accurately.
Study on neural network dynamics in high dimensions with quadratic activation.
Improved machine learning models outperform their simpler counterparts by using imperfect labels.
Student performance prediction - where a machine forecasts the future performance of students as they interact with online coursework - is a challenging problem. Reliable early-stage predictions of a student's future performance could be critical to facilitate timely educational interventions during a course. However, …
A new training method improves MLIPs for faster, lighter simulations.
Contributions: Prior studies on education have mostly followed the model of the cross sectional study, namely, examining the pretest and the posttest scores. This paper shows that students' knowledge throughout the intervention can be estimated by time series analysis using a hidden Markov model. Background: Analyzing …
Unified framework for generating meteorological time series from text.
Proposes a new Bayesian mixture of student-t processes for modeling non-stationary data.
GRASPEL learns large graphs from data efficiently.
C-VAE improves class representation in long-tailed generative models.
The study uses Hidden Markov Models to analyze student enrollment patterns and academic performance.
An important, yet largely unstudied, problem in student data analysis is to detect misconceptions from students' responses to open-response questions. Misconception detection enables instructors to deliver more targeted feedback on the misconceptions exhibited by many students in their class, thus improving the quality…
We analyze the Hessian spectra of large models up to 100B parameters.
Compared to machines, humans are extremely good at classifying images into categories, especially when they possess prior knowledge of the categories at hand. If this prior information is not available, supervision in the form of teaching images is required. To learn categories more quickly, people should see important…
Paper develops methods for estimating and simulating a Student-t Lévy regression model.
Automatic analysis of teacher and student interactions could be very important to improve the quality of teaching and student engagement. However, despite some recent progress in utilizing multimodal data for teaching and learning analytics, a thorough analysis of a rich multimodal dataset coming for a complex real lea…
An important form of prior information in clustering comes in form of cannot-link and must-link constraints. We present a generalization of the popular spectral clustering technique which integrates such constraints. Motivated by the recently proposed -spectral clustering for the unconstrained problem, our method is…
Additive Bayesian networks are types of graphical models that extend the usual Bayesian generalized linear model to multiple dependent variables through the factorisation of the joint probability distribution of the underlying variables. When fitting an ABN model, the choice of the prior of the parameters is of crucial…
Bayesian model identifies skill difficulties and student subgroups in engineering education.
Gaussian process priors are commonly used in aerospace design for performing Bayesian optimization. Nonetheless, Gaussian processes suffer two significant drawbacks: outliers are a priori assumed unlikely, and the posterior variance conditioned on observed data depends only on the locations of those data, not the assoc…
Spectral dimensionality reduction methods enable linear separations of complex data with high-dimensional features in a reduced space. However, these methods do not always give the desired results due to irregularities or uncertainties of the data. Thus, we consider aggressively modifying the scales of the features to …
MCMC complexity matches optimization for large and .
In this paper we do the first large scale analysis of writing style development among Danish high school students. More than 10K students with more than 100K essays are analyzed. Writing style itself is often studied in the natural language processing community, but usually with the goal of verifying authorship, assess…
Improved VAE for heavy-tailed data using Student's t-distributions.
Large scale machine learning (ML) systems such as the Alexa automatic speech recognition (ASR) system continue to improve with increasing amounts of manually transcribed training data. Instead of scaling manual transcription to impractical levels, we utilize semi-supervised learning (SSL) to learn acoustic models (AM) …
Enhances model compression with multi-teacher knowledge distillation.
New diffusion models capture heavy-tailed distributions better.
In this paper, we develop a Bayesian evidence maximization framework to solve the sparse non-negative least squares (S-NNLS) problem. We introduce a family of probability densities referred to as the Rectified Gaussian Scale Mixture (R- GSM) to model the sparsity enforcing prior distribution for the solution. The R-GSM…
New method controls linear systems with adversarial disturbances.
Analyzes Hessian spectrum for neural networks near optimal learning.
A framework predicts employment status for students considering unconscious biases.
We study heterogeneity in the effect of a mindset intervention on student-level performance through an observational dataset from the National Study of Learning Mindsets (NSLM). Our analysis uses machine learning (ML) to address the following associated problems: assessing treatment group overlap and covariate balance,…