Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

4693139185 · Jun 202019922001200920172026
48 results for halving strategy

Seesaw optimizes training by balancing learning rate and batch size, accelerating model pretraining.

problem Optimizing training efficiency for large language models with adaptive optimizers.
method Develops a principled framework for batch-size scheduling, introducing Seesaw which multiplies learning rate by 1/√2 and doubles batch size.
result Empirically, Seesaw reduces wall-clock time by approximately 36% compared to cosine decay, matching theoretical limits.

The medoid of a set of n points is the point in the set that minimizes the sum of distances to other points. It can be determined exactly in O(n^2) time by computing the distances between all pairs of points. Previous works show that one can significantly reduce the number of distance computations needed by adaptively …

2019-06-11abs ↗pdf ↗

This paper proposes a new AED framework for multi-metric experiments with fixed budget.

problem Statistical power challenges in testing multiple metrics simultaneously.
method Two-phase structure: adaptive exploration followed by validation. SHRVar algorithm with relative-variance-based sampling.
result Achieves provable error probability that decreases exponentially.

Derives a size premium from automated market makers in decentralized AI subnets.

problem Determining the profitability and risk of decentralized AI subnets.
method Analyzes daily data on 128 subnets, tests the size premium, and calculates transaction costs.
result The size premium is reduced by a halving of token emissions but remains profitable only below a certain asset threshold.

In this article, we show the existence of conjugations on many simply-connected spin 6-manifolds with free integral cohomology. In a certain class the only condition on X^6 to admit a conjugation with fixed point set M^3 is the obvious one: the existence of a degree-halving ring isomorphism between the Z_2-cohomologies…

2010-01-06abs ↗pdf ↗

Bayesian optimization outperforms other methods in hyperparameter tuning for reinforcement learning.

problem Finding optimal hyperparameters that generalize across random seeds in reinforcement learning.
method Benchmarked Successive Halving, Random Search, and Bayesian Optimization with and without repetitions on PPO2 algorithms for Cartpole and Inverted Pendulum tasks.
result Bayesian optimization with noise robust acquisition function is the best choice.

It is shown that the determinant line bundle associated to a family of Dirac operators over a closed partitioned manifold has a canonical Hermitian metric with compatible connection whose curvature satisfies an additivity formula with contributions from the families of Dirac operators over the two halves. This curvatur…

1998-12-21abs ↗pdf ↗

In their article "The shape of hyperbolic Dehn surgery space," Hodgson and Kerckhoff proved a powerful theorem, half of which they used to make Thurston's Dehn surgery theorem effective. The calculations derived here use both halves of Hodgson and Kerckhoff's theorem to give bounds leading towards a practical algorithm…

2015-04-07abs ↗pdf ↗

To reduce the label complexity in Agnostic Active Learning (A^2 algorithm), volume-splitting splits the hypothesis edges to reduce the Vapnik-Chervonenkis (VC) dimension in version space. However, the effectiveness of volume-splitting critically depends on the initial hypothesis and this problem is also known as target…

2018-09-28abs ↗pdf ↗

We construct higher genus Riemann's minimal surfaces properly embedded in the Euclidean space. To do that we glue end by end a Costa-Hoffman-Meeks examples to two halves genus zero Riemann's minimal surfaces. In first we need to perform a deformation of a Costa-Hoffman-Meeks example to prescribe the flux vector along t…

2005-11-17abs ↗pdf ↗

In this short paper we investigate whether meta-learning techniques can be used to more effectively tune the hyperparameters of machine learning models using successive halving (SH). We propose a novel variant of the SH algorithm (MeSH), that uses meta-regressors to determine which candidate configurations should be el…

2019-09-16abs ↗pdf ↗

Designs efficient factorial experiments for product design under budget constraints.

problem Designing effective experiments for product design with limited traffic and overlapping experiments.
method Two-stage design: first stage samples and infers performance, second stage selects a final policy.
result The method outperforms one-shot tensor completion and unstructured best-arm benchmarks.

Motivated by recent advance of machine learning using Deep Reinforcement Learning this paper proposes a modified architecture that produces more robust agents and speeds up the training process. Our architecture is based on Asynchronous Advantage Actor-Critic (A3C) algorithm where the total input dimensionality is halv…

2018-04-13abs ↗pdf ↗

There are two halves to RL systems: experience collection time and policy learning time. For a large number of samples in rollouts, experience collection time is the major bottleneck. Thus, it is necessary to speed up the rollout generation time with multi-process architecture support. Our work, dubbed WALL-E, utilizes…

2019-01-18abs ↗pdf ↗

Active Learning (AL) is a learning task that requires learners interactively query the labels of the sampled unlabeled instances to minimize the training outputs with human supervisions. In theoretical study, learners approximate the version space which covers all possible classification hypothesis into a bounded conve…

2018-07-24abs ↗pdf ↗

Learning can be seen as approximating an unknown function by interpolating the training data. Kriging offers a solution to this problem based on the prior specification of a kernel. We explore a numerical approximation approach to kernel selection/construction based on the simple premise that a kernel must be good if t…

2018-08-13abs ↗pdf ↗

We derive an optimal policy for adaptively restarting a randomized algorithm, based on observed features of the run-so-far, so as to minimize the expected time required for the algorithm to successfully terminate. Given a suitable Bayesian prior, this result can be used to select the optimal black-box optimization algo…

2019-02-21abs ↗pdf ↗

We introduce a general-purpose conditioning method for neural networks called FiLM: Feature-wise Linear Modulation. FiLM layers influence neural network computation via a simple, feature-wise affine transformation based on conditioning information. We show that FiLM layers are highly effective for visual reasoning - an…

2017-09-22abs ↗pdf ↗

The study bounds the effective diameter of graphs with positive Ollivier curvature.

problem Bounding the effective diameter of graphs with positive Ollivier curvature.
method Introducing reflective graphs and proving discrete Bonnet Myers theorem.
result The effective diameter bound is attained only for specific graphs.

We compute the integer cohomology rings of the ``polygon spaces'' introduced in [Hausmann,Klyachko,Kapovich-Millson]. This is done by embedding them in certain toric varieties; the restriction map on cohomology is surjective and we calculate its kernel using ideas from the theory of Gröbner bases. Since we do not inver…

1997-06-01abs ↗pdf ↗

In earlier studies, the estimation of the volatility of a stock using information on the daily opening, closing, high and low prices has been developed; the additional information in the high and low prices can be incorporated to produce unbiased (or near-unbiased) estimators with substantially lower variance than the …

2008-04-01abs ↗pdf ↗

The two main issues for managing wrong way risk (WWR) for the credit valuation adjustment (CVA, i.e. WW-CVA) are calibration and hedging. Hence we start from a novel model-free worst-case approach based on static hedging of counterparty exposure with liquid options. We say "start from" because we demonstrate that a nai…

2016-09-03abs ↗pdf ↗

This thesis is about the study of Lie groupoids endowed with a compatible (multiplicative) differential 1-form. The motivation and scope of the present work is to study the geometry of PDEs using the formalism of Lie groupoids and multiplicative forms; as such, ideas from the two theories have to be introduced and expl…

2013-06-05abs ↗pdf ↗

The interdependent nature of the global economy has become stronger with increases in international trade and investment. We propose a new model to reconstruct the international trade network and associated cost network by maximizing entropy based on local information about inward and outward trade. We show that the tr…

2018-06-02abs ↗pdf ↗

Bob predicts a future observation based on a sample of size one. Alice can draw a sample of any size before issuing her prediction. How much better can she do than Bob? Perhaps surprisingly, under a large class of loss functions, which we refer to as the Cover-Hart family, the best Alice can do is to halve Bob's risk. …

2012-06-15abs ↗pdf ↗

A particular Riemannian metric which originally has been obtained for a well-known coordinate system in the Euclidean 3-space, is shown to specify, in fact, a manifold with boundary. There are two ways to make the manifold complete. One is to identify two halves of the boundary that turns the manifold into Euclidean 3-…

2005-01-11abs ↗pdf ↗

We fully describe the horofunction boundary hL2\partial_h L_2 with the word metric associated with the generating set {t,at}\{t,at\} (i.e the metric arising in the Diestel-Leader graph DL(2,2)\text{DL}(2,2)). The visual boundary L2\partial_\infty L_2 with this metric is a subset of hL2\partial_h L_2. Although $\partial_\infty L_2…

2014-10-31abs ↗pdf ↗

New ff-vectors reveal geometric Lefschetz-like decompositions of flag spheres.

problem Understanding ff-vectors of balanced simplicial complexes and flag spheres.
method Analyzing hh-vectors and ff-vectors of flag spheres and balanced simplicial complexes.
result Found ff-vectors leading to geometric Lefschetz-like decompositions.

Though machine learning algorithms excel at minimizing the average loss over a population, this might lead to large discrepancies between the losses across groups within the population. To capture this inequality, we introduce and study a notion we call maximum weighted loss discrepancy (MWLD), the maximum (weighted) d…

2019-06-08abs ↗pdf ↗

When applied to training deep neural networks, stochastic gradient descent (SGD) often incurs steady progression phases, interrupted by catastrophic episodes in which loss and gradient norm explode. A possible mitigation of such events is to slow down the learning process. This paper presents a novel approach to contro…

2017-09-05abs ↗pdf ↗