Data science enhances knot theory by analyzing invariant relations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
ALPODS AI diagnoses high-dimensional biomedical data with human-understandable explanations.
Software helps teach latent variable methods in multivariate data analytics.
In an attempt to gather a deeper understanding of how convolutional neural networks (CNNs) reason about human-understandable concepts, we present a method to infer labeled concept data from hidden layer activations and interpret the concepts through a shallow decision tree. The decision tree can provide information abo…
Study identifies negative data externalities affecting model performance on specific groups.
Dependent MMD coresets help compare multiple related datasets.
This work proposes xGEMs or manifold guided exemplars, a framework to understand black-box classifier behavior by exploring the landscape of the underlying data manifold as data points cross decision boundaries. To do so, we train an unsupervised implicit generative model -- treated as a proxy to the data manifold. We …
Many modern Artificial Intelligence (AI) systems make use of data embeddings, particularly in the domain of Natural Language Processing (NLP). These embeddings are learnt from data that has been gathered "from the wild" and have been found to contain unwanted biases. In this paper we make three contributions towards me…
This is a lecture note for the course DS-GA 3001 <Natural Language Understanding with Distributed Representation> at the Center for Data Science , New York University in Fall, 2015. As the name of the course suggests, this lecture note introduces readers to a neural network based approach to natural language understand…
Linear models can grok without understanding, improving generalization.
Develops a logifold structure for understanding datasets.
We perform topological data analysis on the internal states of convolutional deep neural networks to develop an understanding of the computations that they perform. We apply this understanding to modify the computations so as to (a) speed up computations and (b) improve generalization from one data set of digits to ano…
Gen AI improves document understanding but not data analysis in public sector tasks.
Then detection and identification of extreme weather events in large-scale climate simulations is an important problem for risk management, informing governmental policy decisions and advancing our basic understanding of the climate system. Recent work has shown that fully supervised convolutional neural networks (CNNs…
Study finds transparency and model performance metrics increase trust in AutoML systems.
This paper proposes a new method for an optimized mapping of temporal variables, describing a temporal stream data, into the recently proposed NeuCube spiking neural network architecture. This optimized mapping extends the use of the NeuCube, which was initially designed for spatiotemporal brain data, to work on arbitr…
Machine learning (ML) algorithms and machine learning based software systems implicitly or explicitly involve complex flow of information between various entities such as training data, feature space, validation set and results. Understanding the statistical distribution of such information and how they flow from one e…
Interpole learns transparent decision-making policies from data.
Paper improves natural language understanding with less data using a new training method.
In many recent applications, data is plentiful. By now, we have a rather clear understanding of how more data can be used to improve the accuracy of learning algorithms. Recently, there has been a growing interest in understanding how more data can be leveraged to reduce the required training runtime. In this paper, we…
This paper provides a guide to feature importance methods for better scientific inference.
This review covers learning under concept drift, including detection, understanding, and adaptation.
Transformers learn topic structure through embedding and attention mechanisms.
FinALBERT predicts stock prices using labelled Stocktwits data.
NNs accurately predict energy eigenvalues and other physical phenomena in 1D quantum mechanics.
New methods improve brain data analysis from fMRI datasets.
Simplified equation predicts model sensitivity to data.
Quantum models generalize well with little data, challenging traditional generalization theories.
This paper explains how transformers learn from unstructured data in ICL.
ZeroSCROLLS benchmarks zero-shot natural language understanding over long texts.
Meta learning works well with overparameterized models, a phenomenon called 'benign overfitting'.
CERT improves language understanding by contrastively learning sentence-level semantics.
Model integrates multi-view temporal data for better understanding of latent dynamics.
New theory explains how equivariant self-supervised learning improves feature extraction.
Research aims to bridge statistical learning to causal models in AI.
MILCCI integrates labels across categories for better understanding of multi-trial data.
Although deep learning techniques have been successfully applied to many tasks, interpreting deep neural network models is still a big challenge to us. Recently, many works have been done on visualizing and analyzing the mechanism of deep neural networks in the areas of image processing and natural language processing.…
The increasing accessibility of data provides substantial opportunities for understanding user behaviors. Unearthing anomalies in user behaviors is of particular importance as it helps signal harmful incidents such as network intrusions, terrorist activities, and financial frauds. Many visual analytics methods have bee…
With online calendar services gaining popularity worldwide, calendar data has become one of the richest context sources for understanding human behavior. However, event scheduling is still time-consuming even with the development of online calendars. Although machine learning based event scheduling models have automate…
Interactive model analysis, the process of understanding, diagnosing, and refining a machine learning model with the help of interactive visualization, is very important for users to efficiently solve real-world artificial intelligence and data mining problems. Dramatic advances in big data analytics has led to a wide …
Graph machine learning lacks a balanced theory, focusing on expressive power and optimization.
Generalization performance of classifiers in deep learning has recently become a subject of intense study. Deep models, typically over-parametrized, tend to fit the training data exactly. Despite this "overfitting", they perform well on test data, a phenomenon not yet fully understood. The first point of our paper is t…
Understanding learning materials (e.g. test questions) is a crucial issue in online learning systems, which can promote many applications in education domain. Unfortunately, many supervised approaches suffer from the problem of scarce human labeled data, whereas abundant unlabeled resources are highly underutilized. To…
Archetypal analysis helps understand binary data sets.
Improved financial sentiment analysis using simple instruction tuning of LLMs.
Advances understanding of neural network generalization via tensor analysis.
In this paper, we address customer review understanding problems by using supervised machine learning approaches, in order to achieve a fully automatic review aspects categorisation and sentiment analysis. In general, such supervised learning algorithms require domain-specific expert knowledge for generating high quali…
Human-in-the-loop data analysis applications necessitate greater transparency in machine learning models for experts to understand and trust their decisions. To this end, we propose a visual analytics workflow to help data scientists and domain experts explore, diagnose, and understand the decisions made by a binary cl…