Non-experts have long made important contributions to machine learning (ML) by contributing training data, and recent work has shown that non-experts can also help with feature engineering by suggesting novel predictive features. However, non-experts have only contributed features to prediction tasks already posed by e…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Developing active inference agents for edge devices with limited resources.
Introduces hierarchical hyperbolic spaces for non-experts.
The recent successes of deep learning have led to a wave of interest from non-experts. Gaining an understanding of this technology, however, is difficult. While the theory is important, it is also helpful for novices to develop an intuitive feel for the effect of different hyperparameters and structural variations. We …
New interface explains contextual bandits to non-experts.
A framework for faster, better infographic design by non-experts and experts alike.
The paper proposes methods to identify and sample from mixtures of Mallows models for top-k rankings.
We survey Mirzakhani's work relating to Riemann surfaces, which spans about 20 papers. We target the discussion at a broad audience of non-experts.
The last decade has seen huge progress in the development of advanced machine learning models; however, those models are powerless unless human users can interpret them. Here we show how the mind's construction of concepts and meaning can be used to create more interpretable machine learning models. By proposing a nove…
This article is based on the lectures in the Winter Braids V (Pau, Feb. 2015). Main puposel of this is to explain how to compute twisted Alexander polynomials for non-experts.
Deep reinforcement learning has achieved great successes in recent years, however, one main challenge is the sample inefficiency. In this paper, we focus on how to use action guidance by means of a non-expert demonstrator to improve sample efficiency in a domain with sparse, delayed, and possibly deceptive rewards: the…
The present paper are the notes of a mini-course addressed mainly to non-experts. It purpose it to provide a first approach to the theory of mapping class groups of non-orientable surfaces.
ML4Chem offers a user-friendly platform for developing and deploying machine learning models in chemistry.
VR methods improve SGD for faster machine learning.
This article is a survey on the braid groups, the Artin groups, and the Garside groups. It is a presentation, accessible to non-experts, of various topological and algebraic aspects of these groups. It is also a report on three points of the theory: the faithful linear representations, the cohomology, and the geometric…
BIOMRC dataset improves MRC performance, especially for non-experts.
Deep reinforcement learning (deep RL) has achieved superior performance in complex sequential tasks by using deep neural networks as function approximators to learn directly from raw input images. However, learning directly from raw images is data inefficient. The agent must learn feature representation of complex stat…
Topological Data Analysis is a recent and fast growing field providing a set of new topological and geometric tools to infer relevant features for possibly complex data. This paper is a brief introduction, through a few selected topics, to basic fundamental and practical aspects of \tda\ for non experts.
Archetypal analysis approximates data by means of mixtures of actual extreme cases (archetypoids) or archetypes, which are a convex combination of cases in the data set. Archetypes lie on the boundary of the convex hull. This makes the analysis very sensitive to outliers. A robust methodology by means of M-estimators f…
The random forest algorithm, proposed by L. Breiman in 2001, has been extremely successful as a general-purpose classification and regression method. The approach, which combines several randomized decision trees and aggregates their predictions by averaging, has shown excellent performance in settings where the number…
We present a self-contained proof of the Gauss-Bonnet theorem for two-dimensional surfaces embedded in using just classical vector calculus. The exposition should be accessible to advanced undergraduate and non-expert graduate students. It may be viewed as an illustration and exercise in multivariate calculus and…
Paper explains why small-loss criterion works for learning from noisy labels.
Recurrent Neural Networks (RNN) have become competitive forecasting methods, as most notably shown in the winning method of the recent M4 competition. However, established statistical models such as ETS and ARIMA gain their popularity not only from their high accuracy, but they are also suitable for non-expert users as…
This article is a survey article on geometric group theory from the point of view of a non-expert who likes geometric group theory and uses it in his own research. The sections are: classical examples, basics about quasiisometry,properties and invariants of groups invariant under quasiisometry, rigidity, hyperbolic spa…
This is a survey paper focusing on the interplay between the curvature and topology of a Riemannian manifold. The first part of the paper provides a background discussion, aimed at non-experts, of Hopf's pinching problem and the Sphere Theorem. In the second part, we sketch the proof of the Differentiable Sphere Theore…
Concept Hierarchies and Formal Concept Analysis are theoretically well grounded and largely experimented methods. They rely on line diagrams called Galois lattices for visualizing and analysing object-attribute sets. Galois lattices are visually seducing and conceptually rich for experts. However they present important…
Algorithm improves learning by integrating diverse agents' behaviors.
Deep feature fusion improves mitosis counting accuracy.
A popular method for selecting the number of clusters is based on stability arguments: one chooses the number of clusters such that the corresponding clustering results are "most stable". In recent years, a series of papers has analyzed the behavior of this method from a theoretical point of view. However, the results …
Recent success in deep learning has generated immense interest among practitioners and students, inspiring many to learn about this new technology. While visual and interactive approaches have been successfully developed to help people more easily learn deep learning, most existing tools focus on simpler models. In thi…
We will simplify the earlier proofs of Perelman's collapsing theorem of 3-manifolds given by Shioya-Yamaguchi and Morgan-Tian. Among other things, we use Perelman's semi-convex analysis of distance functions to construct the desired local Seifert fibration structure on collapsed 3-manifolds. The verification of Perelma…
Named-entity recognition (NER) aims at identifying entities of interest in a text. Artificial neural networks (ANNs) have recently been shown to outperform existing NER systems. However, ANNs remain challenging to use for non-expert users. In this paper, we present NeuroNER, an easy-to-use named-entity recognition tool…
The Infinite Relational Model (IRM) is a probabilistic model for relational data clustering that partitions objects into clusters based on observed relationships. This paper presents Averaged CVB (ACVB) solutions for IRM, convergence-guaranteed and practically useful fast Collapsed Variational Bayes (CVB) inferences. W…
Axioms of Lie algebroid are discussed in order to review some known aspects for non-experts. In particular, it is shown that a Lie QD-algebroid (i.e. a Lie algebra bracket on the Functions(M)-module F of sections of a vector bundle E over a manifold M which satisfies [X,fY]=f[X,Y]+A(X,f)Y for all X,Y from F, all f from…
We will simplify earlier proofs of Perelman's collapsing theorem for 3-manifolds given by Shioya-Yamaguchi and Morgan-Tian. Among other things, we use Perelman's critical point theory (e.g., multiple conic singularity theory and his fibration theory) for Alexandrov spaces to construct the desired local Seifert fibratio…
Neural networks are becoming more and more popular for the analysis of physiological time-series. The most successful deep learning systems in this domain combine convolutional and recurrent layers to extract useful features to model temporal relations. Unfortunately, these recurrent models are difficult to tune and op…
Bayesian optimization has emerged as a strong candidate tool for global optimization of functions with expensive evaluation costs. However, due to the dynamic nature of research in Bayesian approaches, and the evolution of computing technology, using Bayesian optimization in a parallel computing environment remains a c…
ChemCrow enhances LLMs for chemistry tasks, automating complex chemical processes.
In many machine learning applications, crowdsourcing has become the primary means for label collection. In this paper, we study the optimal error rate for aggregating labels provided by a set of non-expert workers. Under the classic Dawid-Skene model, we establish matching upper and lower bounds with an exact exponent …
Crowdsourcing has become a popular method for collecting labeled training data. However, in many practical scenarios traditional labeling can be difficult for crowdworkers (for example, if the data is high-dimensional or unintuitive, or the labels are continuous). In this work, we develop a novel model for crowdsourcin…
Algorithm selection and hyperparameter tuning remain two of the most challenging tasks in machine learning. Automated machine learning (AutoML) seeks to automate these tasks to enable widespread use of machine learning by non-experts. This paper introduces OBOE, a collaborative filtering method for time-constrained mod…
Mirzakhani's thesis counts geodesics on hyperbolic surfaces, finding a specific asymptotic formula.
Recent progress in AutoML has lead to state-of-the-art methods (e.g., AutoSKLearn) that can be readily used by non-experts to approach any supervised learning problem. Whereas these methods are quite effective, they are still limited in the sense that they work for tabular (matrix formatted) data only. This paper descr…
Object detection is a computer vision field that has applications in several contexts ranging from biomedicine and agriculture to security. In the last years, several deep learning techniques have greatly improved object detection models. Among those techniques, we can highlight the YOLO approach, that allows the const…
Study knot groups from disc patterns in 3D space.
Robo-advisors estimate clients' risk aversion using interactive questionnaires.
For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems. In this work, we explore goals defined in terms of (non-expert) human preferences between pairs of trajectory segments. We show that this approach can effective…
Online health communities are a valuable source of information for patients and physicians. However, such user-generated resources are often plagued by inaccuracies and misinformation. In this work we propose a method for automatically establishing the credibility of user-generated medical statements and the trustworth…