Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

3978116155 · Jun 202019922001200920172026
48 results for sparse conjunctions

Study on list learning with noisy data, showing limits and some learnable cases.

problem Learning from noisy data in a list learning context.
method Inspired by coding theory, extends list learning model to study sparse conjunctions and parities/majors.
result Sparse conjunctions can be efficiently list learned under certain conditions, but parities and majors cannot be efficiently learned.

Bayesian methods improve drug discovery experiment design.

problem Optimizing drug screening experiments in high-dimensional data.
method Bayesian inference and optimisation with upper confidence bound algorithms, Thompson sampling, and sparse tree search.
result Sparse tree search techniques outperform other methods in drug toxicity screening.

Applying machine learning techniques to the quickly growing data in science and industry requires highly-scalable algorithms. Large datasets are most commonly processed "data parallel" distributed across many nodes. Each node's contribution to the overall gradient is summed using a global allreduce. This allreduce is t…

2018-02-22abs ↗pdf ↗

A new algorithm solves sparse optimization problems on measures efficiently.

problem Sparse optimization problems on measures.
method Over-parameterized Stochastic Gradient Descent with Random Features.
result Global convergence with rate O(log(K)/K)O(\log(K)/\sqrt{K}) and bounded total variation norms.

State-of-the-art methods for Convolutional Sparse Coding usually employ Fourier-domain solvers in order to speed up the convolution operators. However, this approach is not without shortcomings. For example, Fourier-domain representations implicitly assume circular boundary conditions and make it hard to fully exploit …

2019-08-31abs ↗pdf ↗

New distributions allow greedy arm selection in sparse bandit problems.

problem Sparse contextual bandit problem with sparse parameters and feature distributions.
method Introduced new distribution classes and demonstrated that mixtures of these distributions are also greedy-applicable.
result Greedy algorithm applicable to a wider range of arm feature distributions, including those with origin-asymmetric support.

As a contribution to interpretable machine learning research, we develop a novel optimization framework for learning accurate and sparse two-level Boolean rules. We consider rules in both conjunctive normal form (AND-of-ORs) and disjunctive normal form (OR-of-ANDs). A principled objective function is proposed to trade …

2016-06-18abs ↗pdf ↗

We consider the problem of designing a sparse Gaussian process classifier (SGPC) that generalizes well. Viewing SGPC design as constructing an additive model like in boosting, we present an efficient and effective SGPC design method to perform a stage-wise optimization of a predictive loss function. We introduce new me…

2012-06-26abs ↗pdf ↗

Many emerging use cases of data mining and machine learning operate on large datasets with data from heterogeneous sources, specifically with both sparse and dense components. For example, dense deep neural network embedding vectors are often used in conjunction with sparse textual features to provide high dimensional …

2019-03-20abs ↗pdf ↗

We address the problem of general supervised learning when data can only be accessed through an (indefinite) similarity function between data points. Existing work on learning with indefinite kernels has concentrated solely on binary/multi-class classification problems. We propose a model that is generic enough to hand…

2012-10-22abs ↗pdf ↗

Reinforcement learning (RL) algorithms allow artificial agents to improve their selection of actions to increase rewarding experiences in their environments. Temporal Difference (TD) Learning -- a model-free RL method -- is a leading account of the midbrain dopamine system and the basal ganglia in reinforcement learnin…

2019-09-04abs ↗pdf ↗

SIBRE boosts reinforcement learning convergence by rewarding improvement over past performance.

problem Improving the rate of convergence in reinforcement learning.
method SIBRE is a reward shaping approach that rewards improvement over the agent's own past performance.
result SIBRE converges faster and more stably to the optimal policy compared to baseline RL algorithms.

In this paper, we investigate community detection in networks in the presence of node covariates. In many instances, covariates and networks individually only give a partial view of the cluster structure. One needs to jointly infer the full cluster structure by considering both. In statistics, an emerging body of work …

2016-07-10abs ↗pdf ↗

Space debris warnings follow a predictable pattern, allowing timely satellite maneuvers.

problem Estimating when fresh information about space debris will arrive.
method Statistical learning model of the message arrival process, specifically a Bayesian Poisson process.
result The average prediction error for the next message arrival time is smaller than baseline predictions.

We consider a Black-Scholes market in which a number of stocks and an index are traded. The simplified Capital Asset Pricing Model is the conjunction of the usual Capital Asset Pricing Model, or CAPM, and the statement that the appreciation rate of the index is equal to its squared volatility plus the interest rate. (T…

2011-11-11abs ↗pdf ↗

Modeling air pollutants using data-driven techniques and sparse identification of nonlinear dynamics.

problem Predicting concentrations of air pollutants using hidden physical laws.
method Sparse identification of nonlinear dynamics (SINDy) for parsimonious systems of ordinary differential equations.
result More than half of the critical points are saddle points, indicating system instability.

New framework tackles high-dimensional reliability analysis using surrogate models and active subspaces.

problem High computational cost and curse of dimensionality in reliability analysis of high-dimensional systems.
method Sparse Active Subspace (SAS) algorithm for identifying low-dimensional manifolds and constructing efficient surrogate models.
result Proposed framework significantly improves accuracy and efficiency of reliability analysis compared to existing methods.

Selecting important features in non-linear or kernel spaces is a difficult challenge in both classification and regression problems. When many of the features are irrelevant, kernel methods such as the support vector machine and kernel ridge regression can sometimes perform poorly. We propose weighting the features wit…

2009-06-24abs ↗pdf ↗

Gaussian processes (GPs) are important models in supervised machine learning. Training in Gaussian processes refers to selecting the covariance functions and the associated parameters in order to improve the outcome of predictions, the core of which amounts to evaluating the logarithm of the marginal likelihood (LML) o…

2018-03-28abs ↗pdf ↗

New estimator improves off-policy evaluation for large action spaces.

problem Conventional importance-weighting approaches suffer from excessive variance in off-policy evaluation for large discrete action spaces.
method Proposes OffCEM estimator based on conjunct effect model (CEM), applying importance weighting only to action clusters and using model-based reward estimation for residual effects.
result Proposed estimator is unbiased under local correctness condition, providing substantial improvements in OPE especially with many actions.

We obtain an expression for the curvature of the Lie group SDiffM\cal M and use it to derive Lukatskii's formula for the case where M\cal M is locally Euclidean. We discuss qualitatively some previous findings for SDiffS2S^{2} in conjunction with our result.

1994-03-16abs ↗pdf ↗

This paper introduces a new method to better understand financial market causality.

problem Lack of comprehensive understanding of distributional causality in financial markets.
method Combines piecewise quantile regression with a piecewise linear embedding scheme.
result Uncovered significant tail-tail causal effects and substantial causal asymmetry in cryptocurrency return series.

This paper, to be regularly updated, lists those prime knots with the fewest possible number of crossings for which values of basic knot invariants, such as the unknotting number or the smooth 4-genus, are unknown. This list is being developed in conjunction with "KnotInfo" (www.indiana.edu/~knotinfo), a web-based tabl…

2005-03-07abs ↗pdf ↗

Query2box embeds complex queries as boxes to handle logical operations in large KGs.

problem Handling complex logical queries on large-scale incomplete knowledge graphs.
method Embed KG entities and queries into a vector space as boxes, handling conjunctions as intersections and disjunctions through Disjunctive Normal Form.
result Query2box achieves up to 25% relative improvement over state-of-the-art methods.

We prove the equality of the analytic torsion and the value at zero of a Ruelle dynamical zeta function associated with an acyclic unitarily flat vector bundle on a closed locally symmetric reductive manifold. This solves a conjecture of Fried. This article should be read in conjunction with an earlier paper by Moscovi…

2016-02-01abs ↗pdf ↗

We construct two knot invariants. The first knot invariant is a sum constructed using linking numbers. The second is an invariant of flat knots and is a formal sum of flat knots obtained by smoothing pairs of crossings. This invariant can be used in conjunction with other flat invariants, forming a family of invariants…

2011-09-14abs ↗pdf ↗

We establish continuous maximal regularity results for parabolic differential operators acting on sections of tensor bundles on Riemannian manifolds. As an application, we show that solutions to the Yamabe flow instantaneously regularize and become real analytic in space and time. The regularity result is obtained by i…

2013-09-09abs ↗pdf ↗

Develops a new GLM framework for claims reserving with adaptive estimation.

problem Accurate assessment of claims reserves with dynamic and dependent claim activity.
method Multivariate evolutionary GLM framework with adaptive particle filtering algorithm.
result Adaptive estimation of evolving factors improves claims reserve accuracy.

It is becoming increasingly important to understand the vulnerability of machine learning models to adversarial attacks. In this paper we study the feasibility of robust learning from the perspective of computational learning theory, considering both sample and computational complexity. In particular, our definition of…

2019-09-12abs ↗pdf ↗

For certain problems involving vector fields, it is possible to find an associated imaginary field that, in conjunction with the first, forms a complex field for which the equation can be solved. This result is generalized to arbitrary Clifford algebras, followed by quaternionic vectors as a special case. All results a…

2002-09-28abs ↗pdf ↗

Trivial links are unique up to number of link components, but they can be hard to recognize from arbitrary diagrams. We define a new measure of the complexity of a link embedding, the crumple, and show how this may be used to measure progress toward a trivial embedding. In conjunction with a modified form of arc presen…

2011-10-13abs ↗pdf ↗

Cross-sectional signatures of market panic were recently discussed on daily time scales in [1], extended here to a study of cross-sectional properties of stocks on intra-day time scales. We confirm specific intra-day patterns of dispersion and kurtosis, and find that the correlation across stocks increases in times of …

2010-10-23abs ↗pdf ↗

A new method uses CPD to efficiently model feature interactions in non-sequential data.

problem Efficiently modeling feature interactions in non-sequential data with high computational and memory costs.
method Implicitly represent model parameters as a tensor, factorize into a compact Tensor Train (TT) format, and use Canonical Polyadic (CP) Decomposition for invariance to feature ordering.
result The proposed CP-based predictor outperforms other TN-based predictors on sparse data and matches neural network performance on dense non-sequential tasks.