Non-experts have long made important contributions to machine learning (ML) by contributing training data, and recent work has shown that non-experts can also help with feature engineering by suggesting novel predictive features. However, non-experts have only contributed features to prediction tasks already posed by e…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A framework for faster, better infographic design by non-experts and experts alike.
New interface explains contextual bandits to non-experts.
Introduces hierarchical hyperbolic spaces for non-experts.
Developing active inference agents for edge devices with limited resources.
Selecting an optimal set of icons is a crucial step in the pipeline of visual design to structure and navigate through content. However, designing the icons sets is usually a difficult task for which expert knowledge is required. In this work, to ease the process of icon set selection to the users, we propose a similar…
Robo-advisors estimate clients' risk aversion using interactive questionnaires.
The paper proposes methods to identify and sample from mixtures of Mallows models for top-k rankings.
We survey Mirzakhani's work relating to Riemann surfaces, which spans about 20 papers. We target the discussion at a broad audience of non-experts.
This article is based on the lectures in the Winter Braids V (Pau, Feb. 2015). Main puposel of this is to explain how to compute twisted Alexander polynomials for non-experts.
Deep reinforcement learning has achieved great successes in recent years, however, one main challenge is the sample inefficiency. In this paper, we focus on how to use action guidance by means of a non-expert demonstrator to improve sample efficiency in a domain with sparse, delayed, and possibly deceptive rewards: the…
The present paper are the notes of a mini-course addressed mainly to non-experts. It purpose it to provide a first approach to the theory of mapping class groups of non-orientable surfaces.
Recent progress in AutoML has lead to state-of-the-art methods (e.g., AutoSKLearn) that can be readily used by non-experts to approach any supervised learning problem. Whereas these methods are quite effective, they are still limited in the sense that they work for tabular (matrix formatted) data only. This paper descr…
ML4Chem offers a user-friendly platform for developing and deploying machine learning models in chemistry.
ChemCrow enhances LLMs for chemistry tasks, automating complex chemical processes.
DOCKSTRING simplifies docking simulations for better drug design benchmarks.
This article is a survey on the braid groups, the Artin groups, and the Garside groups. It is a presentation, accessible to non-experts, of various topological and algebraic aspects of these groups. It is also a report on three points of the theory: the faithful linear representations, the cohomology, and the geometric…
BIOMRC dataset improves MRC performance, especially for non-experts.
Deep reinforcement learning (deep RL) has achieved superior performance in complex sequential tasks by using deep neural networks as function approximators to learn directly from raw input images. However, learning directly from raw images is data inefficient. The agent must learn feature representation of complex stat…
Topological Data Analysis is a recent and fast growing field providing a set of new topological and geometric tools to infer relevant features for possibly complex data. This paper is a brief introduction, through a few selected topics, to basic fundamental and practical aspects of \tda\ for non experts.
Archetypal analysis approximates data by means of mixtures of actual extreme cases (archetypoids) or archetypes, which are a convex combination of cases in the data set. Archetypes lie on the boundary of the convex hull. This makes the analysis very sensitive to outliers. A robust methodology by means of M-estimators f…
We present a self-contained proof of the Gauss-Bonnet theorem for two-dimensional surfaces embedded in using just classical vector calculus. The exposition should be accessible to advanced undergraduate and non-expert graduate students. It may be viewed as an illustration and exercise in multivariate calculus and…
This article is a survey article on geometric group theory from the point of view of a non-expert who likes geometric group theory and uses it in his own research. The sections are: classical examples, basics about quasiisometry,properties and invariants of groups invariant under quasiisometry, rigidity, hyperbolic spa…
Recent success in deep learning has generated immense interest among practitioners and students, inspiring many to learn about this new technology. While visual and interactive approaches have been successfully developed to help people more easily learn deep learning, most existing tools focus on simpler models. In thi…
The recent successes of deep learning have led to a wave of interest from non-experts. Gaining an understanding of this technology, however, is difficult. While the theory is important, it is also helpful for novices to develop an intuitive feel for the effect of different hyperparameters and structural variations. We …
This is a survey paper focusing on the interplay between the curvature and topology of a Riemannian manifold. The first part of the paper provides a background discussion, aimed at non-experts, of Hopf's pinching problem and the Sphere Theorem. In the second part, we sketch the proof of the Differentiable Sphere Theore…
Data cleansing is a typical approach used to improve the accuracy of machine learning models, which, however, requires extensive domain knowledge to identify the influential instances that affect the models. In this paper, we propose an algorithm that can suggest influential instances without using any domain knowledge…
Algorithm improves learning by integrating diverse agents' behaviors.
A framework for privacy-preserving DNN pruning and acceleration.
A popular method for selecting the number of clusters is based on stability arguments: one chooses the number of clusters such that the corresponding clustering results are "most stable". In recent years, a series of papers has analyzed the behavior of this method from a theoretical point of view. However, the results …
Method captures fabric mechanics from depth images without expensive setups.
The last decade has seen huge progress in the development of advanced machine learning models; however, those models are powerless unless human users can interpret them. Here we show how the mind's construction of concepts and meaning can be used to create more interpretable machine learning models. By proposing a nove…
We will simplify the earlier proofs of Perelman's collapsing theorem of 3-manifolds given by Shioya-Yamaguchi and Morgan-Tian. Among other things, we use Perelman's semi-convex analysis of distance functions to construct the desired local Seifert fibration structure on collapsed 3-manifolds. The verification of Perelma…
Algorithm selection and hyperparameter tuning remain two of the most challenging tasks in machine learning. Automated machine learning (AutoML) seeks to automate these tasks to enable widespread use of machine learning by non-experts. This paper introduces OBOE, a collaborative filtering method for time-constrained mod…
Named-entity recognition (NER) aims at identifying entities of interest in a text. Artificial neural networks (ANNs) have recently been shown to outperform existing NER systems. However, ANNs remain challenging to use for non-expert users. In this paper, we present NeuroNER, an easy-to-use named-entity recognition tool…
Axioms of Lie algebroid are discussed in order to review some known aspects for non-experts. In particular, it is shown that a Lie QD-algebroid (i.e. a Lie algebra bracket on the Functions(M)-module F of sections of a vector bundle E over a manifold M which satisfies [X,fY]=f[X,Y]+A(X,f)Y for all X,Y from F, all f from…
We will simplify earlier proofs of Perelman's collapsing theorem for 3-manifolds given by Shioya-Yamaguchi and Morgan-Tian. Among other things, we use Perelman's critical point theory (e.g., multiple conic singularity theory and his fibration theory) for Alexandrov spaces to construct the desired local Seifert fibratio…
Bayesian optimization has emerged as a strong candidate tool for global optimization of functions with expensive evaluation costs. However, due to the dynamic nature of research in Bayesian approaches, and the evolution of computing technology, using Bayesian optimization in a parallel computing environment remains a c…
In many machine learning applications, crowdsourcing has become the primary means for label collection. In this paper, we study the optimal error rate for aggregating labels provided by a set of non-expert workers. Under the classic Dawid-Skene model, we establish matching upper and lower bounds with an exact exponent …
Survey on making reinforcement learning models more understandable.
Mirzakhani's thesis counts geodesics on hyperbolic surfaces, finding a specific asymptotic formula.
Object detection is a computer vision field that has applications in several contexts ranging from biomedicine and agriculture to security. In the last years, several deep learning techniques have greatly improved object detection models. Among those techniques, we can highlight the YOLO approach, that allows the const…
The random forest algorithm, proposed by L. Breiman in 2001, has been extremely successful as a general-purpose classification and regression method. The approach, which combines several randomized decision trees and aggregates their predictions by averaging, has shown excellent performance in settings where the number…
Study knot groups from disc patterns in 3D space.
One aim of data mining is the identification of interesting structures in data. For better analytical results, the basic properties of an empirical distribution, such as skewness and eventual clipping, i.e. hard limits in value ranges, need to be assessed. Of particular interest is the question of whether the data orig…
For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems. In this work, we explore goals defined in terms of (non-expert) human preferences between pairs of trajectory segments. We show that this approach can effective…
Online health communities are a valuable source of information for patients and physicians. However, such user-generated resources are often plagued by inaccuracies and misinformation. In this work we propose a method for automatically establishing the credibility of user-generated medical statements and the trustworth…
With super-resolution optical microscopy, it is now possible to observe molecular interactions in living cells. The obtained images have a very high spatial precision but their overall quality can vary a lot depending on the structure of interest and the imaging parameters. Moreover, evaluating this quality is often di…