Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

71142213284 · Jun 202019922001200920172026
48 results for occupancy measures

Study measures gender bias in machine translation using multiple reference points.

problem Measuring and identifying gender bias in machine translation.
method Used an optimal non-biased translator, reference points from occupational statistics and survey.
result Found bias against both genders, but more against women, and found occupations have a greater effect than adjectives.

We present a new method of learning a continuous occupancy field for use in robot navigation. Occupancy grid maps, or variants of, are possibly the most widely used and accepted method of building a map of a robot's environment. Various methods have been developed to learn continuous occupancy maps and have successfull…

2019-10-18abs ↗pdf ↗

Paper uses Sinkhorn distances to improve imitation learning effectiveness.

problem Improving imitation learning algorithms by comparing occupancy measures.
method Formulates imitation learning as Sinkhorn distance minimization, combining optimal transport and cosine distances.
result Proposes a new critic network and transport plan that guide imitation learning.

We introduce cylindrical projections to simulate infinite-dimensional occupation flows of diffusions.

problem Computational intractability of infinite-dimensional occupation flows of diffusions.
method Introduce cylindrical projections to approximate the occupation flow via a finite-dimensional system.
result Strong convergence of cylindrical projections to the initial process with derived rates.

New algorithm reduces dynamic regret for MDPs with unknown transition and adversarial rewards.

problem Episodic linear mixture MDPs with unknown transition and adversarial rewards.
method Combines occupancy-measure-based global optimization and policy-based variance-aware value-targeted regression.
result Achieves near-optimal dynamic regret of O~(dH3K+HK(H+PˉK))\widetilde{\mathcal{O}}(d \sqrt{H^3 K} + \sqrt{HK(H + \bar{P}_K)}).

The recent wave of AI and automation has been argued to differ from previous General Purpose Technologies (GPTs), in that it may lead to rapid change in occupations' underlying task requirements and persistent technological unemployment. In this paper, we apply a novel methodology of dynamic task shares to a large data…

2020-01-28abs ↗pdf ↗

DSAC improves cooperative MARL with general utilities, converging faster than existing methods.

problem Improving cooperation in multi-agent reinforcement learning with nonlinear utilities.
method Decentralized Shadow Reward Actor-Critic (DSAC) that estimates local occupancy measures and derivatives.
result DSAC converges to ε-stationarity in O(1/ε^2.5) steps with high probability, finding globally optimal policies.

In this paper, we obtain analytical expression for the distribution of the occupation time in the red (below level 00) up to an (independent) exponential horizon for spectrally negative Lévy risk processes and refracted spectrally negative Lévy risk processes. This result improves the existing literature in which only…

2019-03-09abs ↗pdf ↗

Digital tools may hinder or facilitate multidisciplinary collaboration in occupational health.

problem Challenges in multidisciplinary collaboration in occupational health.
method Study of a precursor SPSTI's digital tools and their impacts.
result Digital tools can act as both brakes and levers for multidisciplinary dynamics.

A new method estimates SDEs using occupation kernels.

problem Learning multivariate stochastic differential equations (SDEs).
method Two-step procedure: estimate drift, then diffusion. Occupation kernels used in RKHS.
result Validated on simulated and real-world data.

Spacematch matches office workers with suitable workspaces based on their preferences.

problem Matching workers with suitable workspaces in ABW environments.
method Developed a web-based mobile app to collect occupant preferences and IoT sensors for real-time environmental feedback.
result Occupant preferences can be segmented and matched to build a recommendation platform.

We find the optimal investment strategy to minimize the expected time that an individual's wealth stays below zero, the so-called {\it occupation time}. The individual consumes at a constant rate and invests in a Black-Scholes financial market consisting of one riskless and one risky asset, with the risky asset's price…

2008-05-26abs ↗pdf ↗

Investment strategies in occupational pension plans are optimized for non-tradable income risk.

problem Optimizing investment strategies for occupational pension plans in the presence of non-tradable income risk.
method Formulated as a stochastic optimization problem, analyzed in both constant and stochastic volatility environments.
result Random contributions induce the optimal glide path structure, influenced by initial wealth, contributions, and risk aversion.

We consider the problem of robustly maximizing the growth rate of investor wealth in the presence of model uncertainty. Possible models are all those under which the assets' region EE and instantaneous covariation cc are known, and where additionally the assets are stable in that their occupancy time measures converg…

2018-01-19abs ↗pdf ↗

We address the problem of finding an optimal policy in a Markov decision process under a restricted policy class defined by the convex hull of a set of base policies. This problem is of great interest in applications in which a number of reasonably good (or safe) policies are already known and we are only interested in…

2018-02-26abs ↗pdf ↗

Imitation learning seeks to learn an expert policy from sampled demonstrations. However, in the real world, it is often difficult to find a perfect expert and avoiding dangerous behaviors becomes relevant for safety reasons. We present the idea of \textit{learning to avoid}, an objective opposite to imitation learning …

2019-09-24abs ↗pdf ↗

Neural Index Policy for multi-action bandits with heterogeneous budgets.

problem Real-world settings often involve multiple interventions with heterogeneous costs and constraints, breaking classical assumptions.
method Introduces a Neural Index Policy (NIP) that learns to assign budget-aware indices to arm-action pairs using a neural network and differentiable knapsack layer.
result Empirically achieves near-optimal performance while strictly enforcing heterogeneous budgets and scaling to hundreds of arms.

D2SRM solves complex PDEs using deep learning.

problem High-dimensional, Hessian-dependent fully nonlinear parabolic PDEs.
method Single scalar space-time network generating derivative-consistent approximations trained through residuals and penalties.
result Well-posedness and convergence theory established for globally Lipschitz equations.

Occupant behavior (OB) and in particular window openings need to be considered in building performance simulation (BPS), in order to realistically model the indoor climate and energy consumption for heating ventilation and air conditioning (HVAC). However, the proposed OB window opening models are often biased towards …

2018-07-10abs ↗pdf ↗

This work explores efficient reinforcement learning with density features in low-rank MDPs.

problem Efficient reinforcement learning with density features in low-rank MDPs.
method Proposes algorithms for off-policy estimation and online construction of exploratory data distributions.
result Demonstrates sample-efficient learning with density features in low-rank MDPs, overcoming technical challenges.

We provide new bounds on a flux integral over the portion of the boundary of one regular domain contained inside a second regular domain, based on properties of the second domain rather than the first one. This bound is amenable to numerical computation of a flux through the boundary of a domain, for example, when ther…

2013-10-14abs ↗pdf ↗

MOCK learns complex systems from trajectories efficiently.

problem Learning nonparametric differential equations from high-dimensional data.
method MOCK uses multivariate occupation kernel functions to learn vector fields linearly.
result MOCK outperforms other methods on various datasets.

EBIL simplifies IL by estimating expert energy as reward, achieving effective performance.

problem Recovering optimal policy from expert demonstrations without reward signals.
method EBIL uses a two-stage solution: first estimating expert energy as reward, then learning policy.
result EBIL achieves effective performance and interpretable reward signals.

Algorithm speeds up search for stationary targets with guaranteed accuracy.

problem Minimize search time while ensuring high detection accuracy of stationary targets.
method Multi-fidelity Gaussian process model and EMTS algorithm.
result Guaranteed performance in target detection accuracy and search time.

Optimal policy for multi-armed multi-action bandits with unknown parameters.

problem Optimal sequential action selection for multi-armed multi-action bandits with unknown parameters.
method Occupancy-Measured-Reward Index Policy (OMRIP) and R(MA)^2B-UCB algorithm.
result Asymptotically optimal policy with sub-linear regret and low computational complexity.

Study shows visual feedback and monetary incentives reduce plugload energy consumption in commercial buildings.

problem Mitigating energy consumption in commercial buildings through occupant plugload control.
method Field experiments with visual feedback and monetary incentives in government and university buildings.
result Mean energy reduction of ~9.52% in office environments and ~21.61% in university environments with visual feedback.