This research creates and classifies datasets for Setswana and Sepedi news headlines.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Partial-input models fail to detect dataset artifacts, even when they perform poorly.
This paper presents thirteen datasets for binary, multiclass and multilabel classification based on the European Court of Human Rights judgments since its creation. The interest of such datasets is explained through the prism of the researcher, the data scientist, the citizen and the legal practitioner. Contrarily to m…
Improves labeling quality in machine learning with pairwise feedback.
The growth of the modern knowledge-based economy is becoming less and less dependent on tangible assets and more on intangible ones. In this context, the role of human capital in the value creation process has become central. Despite the large amount of scientific work on human capital phenomena, little research has re…
This review is about the convenience, the benefits, as well as the destructive capacities of money. It deals with various aspects of money creation, with its value, and its appropriation. All sorts of money tend to get corrupted by eventually creating too much of them. In the long run, this renders money worthless and …
Study explores cyberbullying datasets and classifier generalization.
ReSkill reconciles RL skill creation with policy optimization.
New optimal transport method handles mass creation and destruction.
Automated test model creation from semi-structured requirements.
Designing AI market for content creation
Proposes a model to optimize feedback for content creators on social media.
Large speech dataset for commercial use with 9.98% word error rate.
We present a method for translating music across musical instruments, genres, and styles. This method is based on a multi-domain wavenet autoencoder, with a shared encoder and a disentangled latent space that is trained end-to-end on waveforms. Employing a diverse training dataset and large net capacity, the domain-ind…
MTS-CycleGAN adapts multivariate time series data for ironmaking industry.
This paper gives an overview of several key innovations in the 19th century which led to complex geometry in the 20th century. This includes the creation of the complex plane, the work of Abel on addition theorems for generalized elliptic integrals, the theory of elliptic functions, holomorphic functions, and the creat…
We study the possibility of completing data bases of a sample of governance, diversification and value creation variables by providing a well adapted method to reconstruct the missing parts in order to obtain a complete sample to be applied for testing the ownership-structure/diversification relationship. It consists o…
We build new test sets for the CIFAR-10 and ImageNet datasets. Both benchmarks have been the focus of intense research for almost a decade, raising the danger of overfitting to excessively re-used test sets. By closely following the original dataset creation processes, we test to what extent current classification mode…
Method for creating synthetic multi-fidelity data sets.
The extremely useful method of Malliavin calculus has not yet gained adequate popularity because of the complicated analytic apparatus of this method. The author attempts here to propose a simplified algebraic formalism similar to Malliavin calculus, but based on the notion of creation-annihilation operators instead of…
Model explains money creation under regulatory constraints.
New handwritten digits dataset for Kannada script.
Entity resolution (ER) is one of the fundamental problems in data integration, where machine learning (ML) based classifiers often provide the state-of-the-art results. Considerable human effort goes into feature engineering and training data creation. In this paper, we investigate a new problem: Given a dataset D_T fo…
Considering the creation of persistence landscape on a parametrized curve and structure of sampling, there exists a random process for which a finite mixture model of persistence landscape (FMMPL) can provide a better description for a given dataset. In this paper, a nonparametric approach for computing integrated mean…
Study shows noisy data collection in ImageNet leads to biased model performance.
Solve Painleve VI to relate instanton bundles.
Accurately predicting customer churn using large scale time-series data is a common problem facing many business domains. The creation of model features across various time windows for training and testing can be particularly challenging due to temporal issues common to time-series data. In this paper, we will explore …
Model generates music to connect missing parts, leveraging latent space of VAE.
Algorithm generates realistic metaorders from public trade data.
HyperStream processes streaming data with workflow creation.
A large volume of research has considered the creation of predictive models for clinical data; however, much existing literature reports results using only a single source of data. In this work, we evaluate the performance of models trained on the publicly-available eICU Collaborative Research Database. We show that cr…
New KD-tree based method for private synthetic data generation.
CLIP dataset helps extract action items from hospital discharge notes.
Graph clustering uses multiscale community detection for improved performance.
FSD50K provides an open dataset of over 51k audio clips for sound event recognition.
A dataset for evaluating engagement with scientific video lectures.
New algorithm optimizes Hölder continuous functions efficiently.
The basic financial purpose of corporation is creation of its value. Liquidity management should also contribute to realization of this fundamental aim. Many of the current asset management models that are found in financial management literature assume book profit maximization as the basic financial purpose. These boo…
The paper proposes a method to evaluate ML models for subjective inference, focusing on sentence toxicity.
Object detection is a computer vision field that has applications in several contexts ranging from biomedicine and agriculture to security. In the last years, several deep learning techniques have greatly improved object detection models. Among those techniques, we can highlight the YOLO approach, that allows the const…
In this demo paper, we introduce the DARPA D3M program for automatic machine learning (ML) and JPL's MARVIN tool that provides an environment to locate, annotate, and execute machine learning primitives for use in ML pipelines. MARVIN is a web-based application and associated back-end interface written in Python that e…
FedMD combines model distillation for federated learning with private models.
Unified analytic account of correlation emergence and Epps effect in coupled limit order books
The Internet has rich and rapidly increasing sources of high quality educational content. Inferring prerequisite relations between educational concepts is required for modern large-scale online educational technology applications such as personalized recommendations and automatic curriculum creation. We present PREREQ,…
Generative adversarial networks have been able to generate striking results in various domains. This generation capability can be general while the networks gain deep understanding regarding the data distribution. In many domains, this data distribution consists of anomalies and normal data, with the anomalies commonly…
In this short article I introduce the knotR package, which creates two dimensional knot diagrams optimized for visual appearance using the R programming language. The knotR package is a systematic R-centric suite of software for the creation of production-quality artwork of knot diagrams, released under GPL2.
System suggests clinical concepts in real-time for faster note creation.
The paper describes a new geometric structure for general Clifford algebras.