Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,932 papers · 148 categories

Trend · papers per month

0.5%1.0%1.5%2.0% · Apr 202019922001200920172026
48 results for COLT 2019

New algorithms reduce overfitting in multiclass classification.

problem Excessive reuse of test datasets in machine learning leads to overfitting, especially in multiclass classification.
method Developed computationally efficient algorithms to reduce overfitting bias in multi-class classification.
result Achieved overfitting bias of Θ(√(k/(mn)), k/n), matching known upper bounds.

Improved private learning of halfspaces with reduced sample complexity.

problem Private learning of halfspaces with reduced sample complexity.
method Iterative algorithm for solving linear feasibility problem, improving state-of-the-art results.
result Sample complexity reduced to d2.52logGd^{2.5} \cdot 2^{\log^*|G|}, improving d2d^2 factor.

CoLT assesses neural posterior estimates by detecting discrepancies across conditioning inputs.

problem Validating neural posterior estimates from limited data.
method Conditional Localization Test (CoLT) learns a localization function to detect strong deviations.
result CoLT provides rigorous guarantees and practical scalability for comparing true and neural posterior distributions.

Quadratic memory is essential for optimal convex optimization queries.

problem Optimal query complexity for convex optimization and feasibility problems.
method Lower bounds on query complexity for convex optimization and feasibility problems.
result Center-of-mass algorithms are Pareto-optimal for both convex optimization and feasibility problems.

This work accelerates gradient descent with anytime convergence guarantees.

problem Improving the convergence rate of gradient descent methods.
method Proposes a stepsize schedule for gradient descent that achieves anytime convergence rates.
result Gradient descent can achieve convergence rates of O(T1.119)O(T^{-1.119}) for any stopping time TT.

The paper solves open questions in computable PAC learning, providing a complete landscape.

problem Understanding the boundaries and capabilities of computable PAC learning.
method Analyzing and constructing decidable hypothesis classes with different sample complexities and Littlestone dimensions.
result A complete understanding of CPAC learnability, answering open questions and confirming conjectures.

Study on list learning with noisy data, showing limits and some learnable cases.

problem Learning from noisy data in a list learning context.
method Inspired by coding theory, extends list learning model to study sparse conjunctions and parities/majors.
result Sparse conjunctions can be efficiently list learned under certain conditions, but parities and majors cannot be efficiently learned.

Novel approach to universal online learning for bounded losses, closing open problems.

problem Characterizing processes for universal online learning under non-i.i.d. conditions.
method Characterization of processes admitting strong and weak universal learning, introduction of optimistically universal learning rule.
result Introduction of a novel 1NN algorithm that is optimistically universal for bounded losses.

This work proves lower bounds on a greedy teaching set construction algorithm.

problem Characterize the best-case teaching dimension of a concept class.
method A greedy algorithm that iteratively adds points to the teaching set to restrict the concept class the most.
result Lower bounds on the performance of the greedy approach for small k, extending up to k ≤ c*d for small constant c.

New sampling and identity-testing methods for mixtures of distributions that don't satisfy approximate tensorization of entropy.

problem Sampling and identity-testing for mixtures of distributions that don't satisfy approximate tensorization of entropy.
method Fast mixing of Glauber dynamics and efficient identity-testers in the coordinate-conditional sampling access model.
result Efficient identity-testers for mixtures of ATE distributions in the coordinate-conditional sampling access model.

A fast spectral algorithm estimates mean of heavy-tailed vectors efficiently.

problem Estimating the mean of heavy-tailed random vectors with optimal error bound.
method Spectral algorithm using eigenvector computations and novel hyperplane connection.
result Achieves optimal sub-gaussian error bound with improved runtime.

Learning linear predictors with the logistic loss---both in stochastic and online settings---is a fundamental task in machine learning and statistics, with direct connections to classification and boosting. Existing "fast rates" for this setting exhibit exponential dependence on the predictor norm, and Hazan et al. (20…

2018-03-25abs ↗pdf ↗

Analyzed US firm data 1970-2019, identifying scale effects and distributional forms.

problem Understanding differences between small and large firms over time.
method Examined all public US firms, used stylized facts and DLN distribution analysis.
result Small firms are systematically different from large firms, with scale-dependent heteroskedasticity.

A new bandit problem where experiments can be interrupted if results are not promising.

problem Interruptible multi-armed bandit problem with a threshold for cumulative reward.
method Formalized survival regret, identified key components (regret and probability of ruin), derived lower bounds and optimal policies.
result No policy can achieve sublinear survival regret, but optimal policies minimize survival regret in a Pareto sense.

Two methods for quantile regression are compared and found to produce tighter intervals.

problem Comparing methods for producing prediction intervals in quantile regression.
method Two recently proposed methods combining conformal inference and quantile regression.
result Romano et al.'s method typically yields tighter prediction intervals in finite samples.

We propose a hypergraph-based active learning scheme which we term HS2HS^2, HS2HS^2 generalizes the previously reported algorithm S2S^2 originally proposed for graph-based active learning with pointwise queries [Dasarathy et al., COLT 2015]. Our HS2HS^2 method can accommodate hypergraph structures and allows one to ask bo…

2018-11-25abs ↗pdf ↗

Investigates the relationship between US money supply and asset indices over 2001-2019.

problem Determining the relationship between US money supply and asset indices growth.
method Information entropy methodology applied to US asset indices (Property, Russell 2000, S&P 500, NASDAQ) over 2001-2019.
result Growth in US broad money supply is the main determinant of US asset indices growth, especially the NASDAQ and Russell 2000.

Task focuses on fact checking in Q&A forums, improving over baseline systems.

problem Fact checking in community Q&A forums to distinguish factual from opinion.
method Two subtasks: distinguishing factual vs. opinion/advice/socializing, predicting answer truthfulness.
result Improved over baseline systems for both subtasks, but not for Subtask B.

The study of projective varieties with nef anticanonical divisors and log terminal singularities.

problem Understanding the structure and properties of projective varieties with specific divisor conditions.
method Analyzing the Albanese map and MRC fibration for klt projective varieties, showing locally constant fibrations and product decompositions.
result Generalization of results for smooth projective varieties to the klt case, including decomposition into rationally connected and projective varieties with trivial canonical divisor.

Study the Mexican stock market's interdependency structure from 2000-2019.

problem Characterize the interdependency structure of the Mexican Stock Exchange.
method Estimate correlation/concentration matrices from different models and compute network theory metrics.
result Visualizations provide a comprehensive overview of the stock market's interdependency structure.

Study optimizes trading strategies in markets with transaction costs and uncertain models.

problem Optimizing trading strategies in markets with transaction costs and model uncertainty.
method Maximizing worst-case expected utility over a class of models on a filtered probability space.
result Existence of optimal trading strategies for general càdlàg price processes and incomplete filtrations.

Improved sampling from non-log-concave distributions with polynomial query complexity.

problem Sampling from distributions with non-log-concave densities efficiently.
method Combining Ornstein-Uhlenbeck process assumptions and polynomial moment conditions.
result Polynomial query complexity improvement over previous methods.

New proof shows faster convergence rate for robust estimation with Lasso in adversarially contaminated outputs.

problem Robust estimation of parameters in the presence of adversarial output contamination.
method Extended Lasso with Huber loss function and L1L_1 penalty, focusing on specific properties of the Huber function.
result Same convergence rate as Dalalyan and Thompson (2019), but with a different proof.

Study improves CNNs for audio scene classification by restricting receptive fields and adding frequency awareness.

problem Improving CNNs for robust acoustic scene classification.
method Investigated different receptive field configurations for various CNN architectures and introduced Frequency Aware CNNs.
result Several well-performing submissions to DCASE 2019 Challenge were achieved.

AVEC 2019 challenges AI in detecting depression and cross-cultural emotions.

problem Detecting depression and cross-cultural emotions from audiovisual data.
method Comparison of machine learning methods under standardized conditions.
result Baseline system performance on state-of-mind, depression, and cross-cultural tasks.