Crowdsourcing has become an effective and popular tool for human-powered computation to label large datasets. Since the workers can be unreliable, it is common in crowdsourcing to assign multiple workers to one task, and to aggregate the labels in order to obtain results of high quality. In this paper, we provide finit…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Crowdsourcing is an effective tool for human-powered computation on many tasks challenging for computers. In this paper, we provide finite-sample exponential bounds on the error rate (in probability and in expectation) of hyperplane binary labeling rules under the Dawid-Skene crowdsourcing model. The bounds can be appl…
We revisit the classical decision-theoretic problem of weighted expert voting from a statistical learning perspective. In particular, we examine the consistency (both asymptotic and finitary) of the optimal Nitzan-Paroush weighted majority and related rules. In the case of known expert competence levels, we give sharp …
New voting rules protect against strategic voting by robust statistics.
No fair and strategy-proof automated market maker exists for more than two assets.
A method for learning rankings in non-stationary data streams.
Novel analysis improves weighted majority vote in multiclass classification.
We tackle the issue of classifier combinations when observations have multiple views. Our method jointly learns view-specific weighted majority vote classifiers (i.e. for each view) over a set of base voters, and a second weighted majority vote classifier over the set of these view-specific weighted majority vote class…
One unexamined assumption in foreign ownership regulation is the notion that majority voting rights translate to 'effective control'. This assumption is so deeply entrenched in foreign investments law that possession of majority voting rights can determine the nationality of a corporation and its capacity to engage in …
For classifying time series, a nearest-neighbor approach is widely used in practice with performance often competitive with or better than more elaborate methods such as neural networks, decision trees, and support vector machines. We develop theoretical justification for the effectiveness of nearest-neighbor-like clas…
A conformal procedure improves CoT reasoning by aggregating reasoning paths and calibrating abstention rules.
Machine learning ensemble improves accuracy by considering minority answers as more likely true.
In many machine learning scenarios, looking for the best classifier that fits a particular dataset can be very costly in terms of time and resources. Moreover, it can require deep knowledge of the specific domain. We propose a new technique which does not require profound expertise in the domain and avoids the commonly…
This paper analyzes voter coalitions in MakerDAO's decentralized governance.
The number of component classifiers chosen for an ensemble greatly impacts the prediction ability. In this paper, we use a geometric framework for a priori determining the ensemble size, which is applicable to most of existing batch and online ensemble classifiers. There are only a limited number of studies on the ense…
We develop ensemble Convolutional Neural Networks (CNNs) to classify the transportation mode of trip data collected as part of a large-scale smartphone travel survey in Montreal, Canada. Our proposed ensemble library is composed of a series of CNN models with different hyper-parameter values and CNN architectures. In o…
Boosting improves accuracy by combining weak learners into a voting classifier.
New bound improves on weighted majority vote risk estimation.
New voting strategies show committee-based consensus can scale efficiently.
RCAM-based ensemble combines binary classifiers using similarity and vote scheme.
This paper solves aggregation of Pareto optimal models by using Bayesian priors and weighted averaging.
Learning compact and interpretable representations is a very natural task, which has not been solved satisfactorily even for simple binary datasets. In this paper, we review various ways of composing experts for binary data and argue that competitive forms of interaction are best suited to learn low-dimensional represe…
Vote-boosting is a sequential ensemble learning method in which the individual classifiers are built on different weighted versions of the training data. To build a new classifier, the weight of each training instance is determined in terms of the degree of disagreement among the current ensemble predictions for that i…
We propose a voted dual averaging method for online classification problems with explicit regularization. This method employs the update rule of the regularized dual averaging (RDA) method, but only on the subsequence of training examples where a classification error is made. We derive a bound on the number of mistakes…
We consider the -ary classification problem via crowdsourcing, where crowd workers respond to simple binary questions and the answers are aggregated via decision fusion. The workers have a reject option to skip answering a question when they do not have the expertise, or when the confidence of answering that questio…
Prefix consistency improves model reliability by weighting answers based on their reproducibility.
A mathematical analysis of the distribution of voting power in the Council of the European Union operating according to the Treaty of Lisbon is presented. We study the effects of Brexit on the voting power of the remaining members, measured by the Penrose--Banzhaf Index. We note that the effects in question are non-mon…
Fast detection of changepoints in linear regression models.
CITE algorithm provides anytime-valid certification of model outputs.
In machine learning, the domain adaptation problem arrives when the test (target) and the train (source) data are generated from different distributions. A key applied issue is thus the design of algorithms able to generalize on a new distribution, for which we have no label information. We focus on learning classifica…
In machine learning, Domain Adaptation (DA) arises when the distribution gen- erating the test (target) data differs from the one generating the learning (source) data. It is well known that DA is an hard task even under strong assumptions, among which the covariate-shift where the source and target distributions diver…
New inequality for ternary variables improves on existing measures.
Design rule check is a critical step in the physical design of integrated circuits to ensure manufacturability. However, it can be done only after a time-consuming detailed routing procedure, which adds drastically to the time of design iterations. With advanced technology nodes, the outcomes of global routing and deta…
We tackle the PAC-Bayesian Domain Adaptation (DA) problem. This arrives when one desires to learn, from a source distribution, a good weighted majority vote (over a set of classifiers) on a different target distribution. In this context, the disagreement between classifiers is known crucial to control. In non-DA superv…
New method improves LLM judge accuracy by accounting for dependencies in aggregated binary labels.
This paper generalizes an important result from the PAC-Bayesian literature for binary classification to the case of ensemble methods for structured outputs. We prove a generic version of the \Cbound, an upper bound over the risk of models expressed as a weighted majority vote that is based on the first and second stat…
Machine learning selects the best prediction rules from noisy data.
Proposes a new k-NN algorithm to improve classification accuracy by removing noise and pseudo-neighbours.
Majority bit estimation in noisy random recursive DAGs.
Ensemble learning is a powerful approach to construct a strong learner from multiple base learners. The most popular way to aggregate an ensemble of classifiers is majority voting, which assigns a sample to the class that most base classifiers vote for. However, improved performance can be obtained by assigning weights…
Two local learning rules are investigated to avoid weight transport in neural networks.
Delegated votes in Uniswap DAO favor parties with less self-owned votes and a16z-affiliated entities.
LoCoV reduces portfolio optimization errors from sample covariance matrices.
Study improves fair opinion aggregation by balancing voter attributes.
A thermodynamic theory explains EU election vote distributions.
Online voting is an emerging feature in social networks, in which users can express their attitudes toward various issues and show their unique interest. Online voting imposes new challenges on recommendation, because the propagation of votings heavily depends on the structure of social networks as well as the content …
We present a Bayesian formulation of weighted stochastic block models that can be used to infer the large-scale modular structure of weighted networks, including their hierarchical organization. Our method is nonparametric, and thus does not require the prior knowledge of the number of groups or other dimensions of the…
New algorithm reduces communication costs in distributed deep learning.