New pruning technique reduces index size for DNNs.
problem Irregular index form in fine-grained pruning limits parallelism and memory usage.
method Proposes a low-rank binary index matrix for efficient compression and decompression.
result Fine-grained pruning with binary matrices achieves lower memory footprint and higher parallelism.
The paper examines the stability of binary choice models using Gini index and scoring indicators.
problem Stability and discriminatory power of binary choice models.
method Derives the real Gini index and incorporates PSI and KS statistics into the model.
result The real Gini index should be less than the calculated Gini index when the population distribution changes.
Study links K-stability of certain surfaces to binary forms, proving stability and non-stability conditions.
problem Investigating K-stability of specific del Pezzo surfaces.
method Relating K-stability to GIT stability of binary forms, proving stability and non-stability conditions.
result K-polystability and non-K-stability of quasi-smooth hypersurfaces.
Cubic predicts stock market indices by fusing stock latent embeddings and converting to binary classification.
problem Challenges in predicting stock market indices due to isolated time series treatment and simple regression.
method Fusion of stock latent embeddings, binary encoding classification, and confidence-guided prediction.
result Cubic outperforms state-of-the-art baselines in stock index prediction tasks.
Paper ranks stocks by compression risk, not volatility.
problem Investment risk not correlated with stock price volatility.
method Binary-ternary compressive coding of price change time series.
result Compression risk is a better indicator of stock investment risk.
Alternative to convolutions using decision trees for neural networks.
problem Replacing complex convolutions with simpler decision-based layers.
method Binary decisions as indices to conditional distributions, trained using backpropagation.
result Performance similar to conventional neural networks, with runtime improvements.
Develops a continuous compliance index for Islamic equity screening.
problem Binary rulebooks lead to inconsistent compliance assessment of firms.
method Integrates six leading financial and business activity standards into a single continuous index.
result Firms with the same pass/fail label can differ significantly in compliance strength.
Novel link classification connects quadratic forms and knot theory.
problem Classifying isotopy classes of links in 3D space.
method Established a correspondence between quadratic forms and isotopy classes of links.
result Class numbers of quadratic number fields measure link distinguishability.
We study the restless bandit associated with an extremely simple scalar Kalman filter model in discrete time. Under certain assumptions, we prove that the problem is indexable in the sense that the Whittle index is a non-decreasing function of the relevant belief state. In spite of the long history of this problem, thi…
Develops a binary tree model for option pricing with skew dynamics.
problem Option pricing in incomplete markets with skew dynamics.
method Binary tree model with skew Brownian motion dynamics.
result Model preserves skewness under both discrete and continuous time limits.
BEGIN network models binary data without parametric assumptions.
problem Conditional independence in non-parametric families of binary data.
method BEGIN network models binary data using sparse linear representations and block factorizations.
result BEGIN network captures conditional independence for arbitrary binary and multinomial variables.
In this paper the Buchen's pricing formulae of (higher order) asset and bond binary options are incorporated into the pricing formula of power binary options and a pricing formula of "the normal distribution standard options" with the maturity payoff related to a power function and the density function of normal distri…
Study the tradeoff between signal distortion and human perception over finite channels.
problem Characterize the distortion-perception tradeoff for finite channels with arbitrary metrics.
method Solve linear programming problems to compute the distortion-perception function and optimal reconstructions.
result DP function is piecewise linear in the perception index.
Retrieving indexed documents, not by their topical content but their writing style opens the door for a number of applications in information retrieval (IR). One application is to retrieve textual content of a certain author X, where the queried IR system is provided beforehand with a set of reference texts of X. Autho…
Transformer pre-training improves stock return prediction accuracy.
problem Improving stock price prediction accuracy for better investment decisions.
method Pre-trained transformer models on TSX index, fine-tuned for individual stocks, compared to LSTM and XGBoost.
result Transformer model achieved lower mean squared error than benchmarks.
A new approach to the understanding of complex behavior of financial markets index using tools from thermodynamics and statistical physics is developed. Physical complexity, a magnitude rooted in Kolmogorov-Chaitin theory is applied to binary sequences built up from real time series of financial markets indexes. The st…
Flexible per-class regularization improves binary classifiers.
problem Improving binary classifiers by addressing outliers and class imbalance.
method Graph-based adaptive regularization with flexible per-class thresholds.
result Flexible thresholds improve classifier performance and address class imbalance.
We obtain an index of the complexity of a random sequence by allowing the role of the measure in classical probability theory to be played by a function we call the generating mechanism. Typically, this generating mechanism will be a finite automata. We generate a set of biased sequences by applying a finite state auto…
Study binary perceptrons' capacity using random duality theory.
problem Characterize the capacity of binary perceptrons with general thresholds.
method Utilized fully lifted random duality theory (fl RDT) to characterize the capacity.
result Characterizations match replica symmetry breaking predictions and uncover the capacity for zero-threshold scenario.
We investigate the effect of tax evasion on the income distribution and the inequality index of a society through a kinetic model described by a set of nonlinear ordinary differential equations. The model allows to compute the global outcome of binary and multiple microscopic interactions between individuals. When evas…
Confidence intervals improve evaluation of binary prediction rules in data mining.
problem Uncertainty in performance measures estimation from finite datasets.
method Asymptotic normal approximations for confidence intervals, with a blurring correction.
result Improved finite sample coverage probabilities and general performance measures inference.
Neural Index Policy for multi-action bandits with heterogeneous budgets.
problem Real-world settings often involve multiple interventions with heterogeneous costs and constraints, breaking classical assumptions.
method Introduces a Neural Index Policy (NIP) that learns to assign budget-aware indices to arm-action pairs using a neural network and differentiable knapsack layer.
result Empirically achieves near-optimal performance while strictly enforcing heterogeneous budgets and scaling to hundreds of arms.
New method optimizes policies without assuming known link functions between preferences and rewards.
problem Policy alignment with unknown and unrestricted link functions.
method Formulates an f-divergence-constrained reward maximization problem, learning policies directly. result Induces a semiparametric single-index binary choice model for policy alignment.
We consider the recovery of regression coefficients, denoted by β0, for a single index model (SIM) relating a binary outcome Y to a set of possibly high dimensional covariates X, based on a large but 'unlabeled' dataset U, with Y never observed. On U, we fully ob…
The paper proposes methods to directly optimize complex classification metrics.
problem Handling class-imbalanced cases with non-decomposable metrics.
method Calibrated surrogate maximization of linear-fractional utility.
result Calibrated surrogate maximization can coincide with true utility maximization under certain conditions.
Extends BBSM model to incorporate ESG ratings and path dynamics.
problem Price stock options considering historical market index dynamics and ESG ratings.
method Develops discrete, binary tree option pricing model under BBSM with ESG valuation.
result Model accurately fits stock price changes and European call option prices.
A new approach to the understanding of the complex behavior of financial markets index using tools from thermodynamics and statistical physics is developed. Physical complexity, a magnitude rooted in the Kolmogorov-Chaitin theory is applied to binary sequences built up from real time series of financial markets indices…
Local polynomial regression (Fan and Gijbels 1996) is an important class of methods for nonparametric density estimation and regression problems. However, straightforward implementation of local polynomial regression has quadratic time complexity which hinders its applicability in large-scale data analysis. In this pap…
The paper discusses thresholds and bounds for accuracy in binary classification systems.
problem The accuracy of binary classification systems and its dependence on prevalence.
method Analyzing the precision-prevalence curve and negative predictive value-prevalence curve to find thresholds and bounds.
result Thresholds (φe and φn) bound various accuracy metrics (Fβ, F1, FM, MCC) and the ratio of maximum accuracy to prevalence. Faster Tsetlin Machines use clause indexing to speed inference and learning.
problem Overfitting and slow inference in Tsetlin Machines.
method Introduced a look-up table that indexes clauses based on feature falsification, enabling faster evaluation of clauses.
result Up to 15 times faster classification and three times faster learning on MNIST and Fashion-MNIST.
The trade-off between the cost of acquiring and processing data, and uncertainty due to a lack of data is fundamental in machine learning. A basic instance of this trade-off is the problem of deciding when to make noisy and costly observations of a discrete-time Gaussian random walk, so as to minimise the posterior var…
New ROC tools assess predictive abilities for any linearly ordered outcomes.
problem Fundamental restriction in ROC analysis for non-dichotomous outcomes.
method ROC movies and UROC curves for linearly ordered outcomes.
result CPA equals AUC for binary outcomes and relates to Spearman's coefficient for pairwise distinct outcomes.
Proposes a quantum-inspired algorithm for selecting representative data subsets.
problem Selecting the most representative subset of data from a larger dataset.
method Uses a Quadratic Unconstrained Binary Optimization (QUBO) problem approach.
result Demonstrates the effectiveness of the selector algorithm in finance applications.
Study reveals heavy-tailed behavior in training ReLU gates.
problem Understanding heavy-tailed distribution in stochastic deep learning.
method Experimental study of heavy-tail index for S.G.D. and a variant.
result Two algorithms exhibit similar heavy-tail behavior on ReLU data.
Binary choice forests model customer choices in retailing.
problem Estimating DCMs using transaction data is challenging and prone to misspecification.
method Random forest of binary decision trees to represent DCMs, interpretable and consistent predictions.
result Random forest can predict choice probabilities and assortments unseen in training data.
New bandit model for healthcare intervention planning.
problem Maximizing patient health with limited monitoring resources.
method Developed Collapsing Bandits model and derived optimal policies.
result 3-order-of-magnitude speedup in algorithm performance.
The paper categorizes and analyzes various event-linked perpetual futures contracts.
problem Developing a risk-design framework for complex event-linked perpetual futures.
method Formal taxonomy of seven pure-form canonical variants, organized along four design axes.
result Detailed analysis of microstructure properties and limitations of various variants.
This project explores several Machine Learning methods to predict movie genres based on plot summaries. Naive Bayes, Word2Vec+XGBoost and Recurrent Neural Networks are used for text classification, while K-binary transformation, rank method and probabilistic classification with learned probability threshold are employe…
We study the dynamical behavior of high-frequency data from the Korean Stock Price Index (KOSPI) using the movement of returns in Korean financial markets. The dynamical behavior for a binarized series of our models is not completely random. The conditional probability is numerically estimated from a return series of K…
Minwise hashing (Minhash) is a widely popular indexing scheme in practice. Minhash is designed for estimating set resemblance and is known to be suboptimal in many applications where the desired measure is set overlap (i.e., inner product between binary vectors) or set containment. Minhash has inherent bias towards sma…
The paper examines bias in ML models using the Adult dataset.
problem Understanding and mitigating bias in machine learning models.
method Mathematical framework for fair learning, Disparate Impact index, and evaluation of bias reduction methods.
result Some common bias reduction methods are ineffective.
Paper introduces new metrics for evaluating model accuracy.
problem Improving model accuracy and calibration.
method Developed two second-order accuracy metrics with integral and numerical representations.
result Validates model calibration settings and reveals distortions.
Develops a new framework for perpetual futures on binary prediction markets.
problem Lack of effective risk management in perpetual futures on binary prediction markets.
method PIRAP framework with six components: index estimator, margin sizing, leverage, funding rule, halt protocol, and eligibility framework.
result Mixed results from empirical evaluation, with some pre-registered floors passing and others failing.
The receiver operating characteristic (ROC) curve is a very useful tool for analyzing the diagnostic/classification power of instruments/classification schemes as long as a binary-scale gold standard is available. When the gold standard is continuous and there is no confirmative threshold, ROC curve becomes less useful…
A scalable algorithm improves AUC optimization for semi-supervised ordinal regression.
problem Optimizing AUC for semi-supervised ordinal regression with limited labeled data.
method Proposes QS3ORAO using quadruply stochastic gradients for scalable kernelized learning. result Converges to optimal solution at O(1/t) rate, demonstrating efficiency and effectiveness. We apply the Zipf power law to financial time series of WIG20 index daily changes (open-close). Thanks to the mapping of time series signal into the sequence of 2k+1 'spin-like' states, where k=0, 1/2, 1, 3/2, ..., we are able to describe any time series increments, with almost arbitrary accuracy, as the one of such 's…
A neural network with a single hidden layer can't represent certain multivariable functions.
problem Representing certain multivariable functions with a neural network having only one hidden layer.
method Developed a continuum version of a one-hidden-layer neural network with ReLU activation, and proved constraints on its parameters and second derivative.
result Existence of a smooth binary function that cannot be precisely represented by any such neural network.
This paper investigates semi-supervised hashing methods using variational autoencoders.
problem Semantic hashing with scarce labels.
method Two semi-supervised approaches: joint modeling and pairwise loss.
result The pairwise approach can improve hash quality with many labeled points but degrades with few labels.