Heuristic tool estimates lactate threshold for easier training decisions.
problem Improving lactate threshold estimation for recreational runners.
method Formalized lactate threshold principles, iterative methodology, heuristic approach.
result Heuristic %60 of 'endurance running speed reserve' is reliable and accessible.
Lactate threshold is considered an essential parameter when assessing performance of elite and recreational runners and prescribing training intensities in endurance sports. However, the measurement of blood lactate concentration requires expensive equipment and the extraction of blood samples, which are inconvenient f…
Typically flat filling, linear or polynomial interpolation methods to generate missing historical data. We introduce a novel optimal method for recreating data generated by a diffusion process. The results are then applied to recreate historical data for stocks.
Model recreates LOB from TAQ data for small-tick stocks.
problem Lack of LOB data makes market analysis difficult.
method Combines GRU and ODE-RNN to predict LOB volumes.
result Model accurately recreates LOB with high precision.
LOBRM model recreates limit order books from trade and quote data.
problem Lack of LOB data and limitations in LOBRM model.
method Extended LOBRM with time-weighted z-score standardization and exponential decay kernel, conducted in chronological order.
result LOBRM with decay kernel outperforms traditional models and module ensembling is effective.
Paper proposes efficient cost functions for automated market makers in DeFi.
problem Inefficient and computationally complex cost functions in DeFi.
method Proposes and analyzes constant circle/ellipse based cost functions.
result Proposed cost functions are computationally efficient and robust against attacks.
Machine learning recreates the periodic table from element properties.
problem Recreating the periodic table using machine learning.
method Unsupervised machine learning with GTM for feature embedding.
result PTG autonomously generates various periodic table layouts.
This is a recreational paper showing that certain linked graphs cannot be separated. The proofs employ elementary covering space theory, an appeal to a theorem of Scharlemann (concerning the band sums of two unknots), and a Jones polynomial calculation.
We recreate an unpublished proof of William Thurston from the early 1970's that any smooth 2-plane field on a manifold of dimension at least 4 is homotopic to the tangent plane field of a foliation.
Some online advertising offers pay only when an ad elicits a response. Randomness and uncertainty about response rates make showing those ads a risky investment for online publishers. Like financial investors, publishers can use portfolio allocation over multiple advertising offers to pursue revenue while controlling r…
A deep learning model is applied for predicting block-level parking occupancy in real time. The model leverages Graph-Convolutional Neural Networks (GCNN) to extract the spatial relations of traffic flow in large-scale networks, and utilizes Recurrent Neural Networks (RNN) with Long-Short Term Memory (LSTM) to capture …
There are several (mathematical) reasons why Dupire's formula fails in the non-diffusion setting. And yet, in practice, ad-hoc preconditioning of the option data works reasonably well. In this note we attempt to explain why. In particular, we propose a regularization procedure of the option data so that Dupire's local …
The study clusters neighborhoods based on childhood vulnerability data and program retention rates.
problem Improving childhood outcomes and decision-making in early childhood development.
method Data-driven approaches combining public and private sources to visualize children's experiences and analyze program retention rates.
result Neighborhoods can be clustered based on childhood vulnerability, and certain programs have higher retention rates.
With the availability of vast amounts of user visitation history on location-based social networks (LBSN), the problem of Point-of-Interest (POI) prediction has been extensively studied. However, much of the research has been conducted solely on voluntary checkin datasets collected from social apps such as Foursquare o…
Dropout training, originally designed for deep neural networks, has been successful on high-dimensional single-layer natural language tasks. This paper proposes a theoretical explanation for this phenomenon: we show that, under a generative Poisson topic model with long documents, dropout training improves the exponent…
Physics-informed neural networks improve baryonic predictions from dark matter simulations.
problem Recreating hydrodynamic simulations from dark matter requires expensive and time-consuming computations.
method Combining neural network architectures with physical constraints and using Kullback-Leibler divergence for prediction comparison.
result Improved accuracy of baryonic predictions based on dark matter halo properties, successful recovery of the metallicity relation, and preserved scatter.
In an informal way, a number of thoughts on the financial crisis 2008 are presented from a physicist's viewpoint, considering the problem as a nonergodicity transition of a spin-glass type of system. Some tentative suggestions concerning the way out of the crisis are also discussed, concerning Keynesian "deficit spendi…
Hybrid model combines deep learning and agent-based methods for synthetic LOB generation.
problem Generating realistic financial time series data for model training.
method Combining TABL model with Chiarella model for intraday trading activity simulation.
result Hybrid model generates realistic price dynamics but fails to accurately recreate market microstructure.
In our model, private actors with interbank cash flows similar to, but nore general than (Carmona, Fouque, Sun, 2013) borrow from the outside economy at a certain interest rate, controlled by the central bank, and invest in risky assets. Each private actor aims to maximize its expected terminal logarithmic utility. The…
Deep learning improves reinforcement learning in video games.
problem Challenges in reinforcement learning with high-dimensional inputs.
method Used deep Q-network and batch normalization to improve reinforcement learning.
result Some agents learned to play T-rex Runner better than human experts.
GANs struggle with discontinuous distributions and object counting in images.
problem GANs' limitations in learning from complex distributions and counting objects.
method Evaluated GANs on synthetic datasets including discontinuous and noisy points, and images with varying polygons.
result GANs fail to accurately recreate discontinuous distributions and count objects in images.
Faster, context-aware topic modeling approach unveiled.
problem Slow and heavy preprocessing in topic modeling.
method Conceptualizes topics as independent axes, decomposes embeddings using ICA.
result 4.5x faster than BERTopic on average, with diverse and coherent topics.
Recent work on imitation learning has generated policies that reproduce expert behavior from multi-modal data. However, past approaches have focused only on recreating a small number of distinct, expert maneuvers, or have relied on supervised learning techniques that produce unstable policies. This work extends InfoGAI…
We identify a trade-off between robustness and accuracy that serves as a guiding principle in the design of defenses against adversarial examples. Although this problem has been widely studied empirically, much remains unknown concerning the theory underlying this trade-off. In this work, we decompose the prediction er…
PAC learning simplified as bipartite matching.
problem Efficiently solving PAC learning problems.
method Transductive learning and one-inclusion graphs.
result PAC learning can be reduced to bipartite matching.
Paper reveals free information from differential privacy mechanisms improving query accuracy.
problem Improving query accuracy with differential privacy mechanisms.
method Analysis of Noisy Max and Sparse Vector mechanisms.
result Noisy Max releases the noisy gap between the approximate maximizer and runner-up.
SASE improves attributed graph clustering for large graphs with linear time and space complexity.
problem Challenges in clustering large attributed graphs due to high computational and memory costs.
method SASE combines node features smoothing, scalable spectral clustering, and adaptive order selection.
result SASE achieves a 6.9% improvement in ACC and a 5.87x speedup on the ArXiv dataset.
A new method quantizes output space for multi-target regression.
problem Predicting multiple continuous targets using shared predictors.
method MRQ method that quantizes output space to model dependencies and scale.
result MRQ achieves high scalability and competitive accuracy.
Enhances graph embeddings by preserving graph topology.
problem Node2vec struggles to recreate the topology of input graphs.
method Introduces a topological loss term to Node2vec, aligning the persistence diagram of the embedding to that of the input graph.
result Reconstructs both geometry and topology of input graphs.
POTATOES improves autoencoder UOD accuracy without tuning.
problem Improving unsupervised outlier detection accuracy.
method Randomly partition data, overfit each part with an autoencoder, use max reconstruction error as anomaly score.
result Significant improvement in UOD performance for dense inlier sets.
Deep learning is having a profound impact in many fields, especially those that involve some form of image processing. Deep neural networks excel in turning an input image into a set of high-level features. On the other hand, tomography deals with the inverse problem of recreating an image from a number of projections.…
Non-negative matrix factorization (NMF) is a dimensionality reduction technique which tends to produce a sparse representation of data. Commonly, the error between the actual and recreated matrices is used as an objective function, but this method may not produce the type of representation we desire as it allows for th…
This paper attempts to define a generalisation of the standard Einstein condition (in conformal/metric geometry) to any parabolic geometry. To do so, it shows that any preserved involution σ of the adjoint bundle $\mc{A}$ gives rise, given certain algebraic conditions, to a unique preferred affine connection ∇…
Random smoothing struggles to certify high-dimensional image robustness.
problem Certifying adversarial robustness for high-dimensional images with p>2. method Analysis of random smoothing for ℓp robustness, focusing on ℓ∞. result Noise distribution required for ℓp robustness must have high variance, leading to trivial classifiers. New method protects whistleblowers from retaliation by ensuring their reports remain private.
problem Whistleblowers face retaliation, and current protections are insufficient.
method Formalizes protection against strong-adversary threat model as per-report (0,δ)-differential privacy, and provides a generic mechanism to reduce private auditing to private continual counting. result Demonstrates a reduction in selection error and improved utility over randomized response.
This paper extends liquidity returns in geometric mean markets to time-varying weights.
problem Understanding returns and no-arbitrage prices in geometric mean markets with time-varying weights.
method Extending known results for constant-weight G3Ms to the general case of G3Ms with time-varying and potentially stochastic weights.
result LP shares can replicate the payoffs of financial derivatives and various trading strategies.
Paper uses K-NN resampling to simulate and evaluate LOB markets.
problem Simulating and evaluating limit order book (LOB) markets.
method Applies K-nearest neighbor (K-NN) resampling to LOB simulation and evaluation. result Demonstrates the effectiveness and efficiency of K-NN resampling in LOB simulation and evaluation. GANs generate realistic cyber-attack alerts with feature dependencies.
problem Challenges in creating realistic cyber-attack alert data.
method Used Generative Adversarial Networks (GANs) to learn complex data distributions.
result GANs successfully generate realistic alerts with feature dependencies.
Random forest model predicts tennis match outcomes with 80% accuracy.
problem Predicting tennis match outcomes before the game starts.
method Used a large database of tennis match information and a random forest model.
result Identified serve strength as a key predictor of match outcome.
Deep neural network classifies DaTscan SPECT images for Parkinson's Disease.
problem Early diagnosis of Parkinson's Disease through objective analysis of SPECT images.
method InceptionV3 architecture with custom binary classifier, 10-fold cross validation.
result Deep neural network achieves high accuracy in classifying DaTscan SPECT images.
Separating a singing voice from its music accompaniment remains an important challenge in the field of music information retrieval. We present a unique neural network approach inspired by a technique that has revolutionized the field of vision: pixel-wise image classification, which we combine with cross entropy loss a…
Bandlimited random neural networks may not approximate all functions perfectly.
problem Expressive power of shallow neural networks with bandlimited random weights.
method Ridgelet analysis for deriving approximation error lower bounds.
result Bandlimited random weights can lead to non-zero approximation error.
Machine learning improves accuracy of running gait event detection from tibial acceleration.
problem Accurate detection of running gait events from tibial acceleration data.
method Structured machine learning models compared to heuristic methods.
result Structured recurrent neural network model offers most accurate estimation of gait events.
Hybrid model simulates market dynamics using neural stochastic background traders.
problem Lack of realistic LOB simulations that combine historical data and dynamic interactions.
method Neural stochastic background trader trained on historical LOB data, embedded in multi-agent simulation.
result Hybrid model recreates stylised market facts and financial herding behaviors.
EarnHFT tackles HFT challenges with hierarchical RL, significantly outperforming existing methods.
problem Challenges in applying RL to HFT due to long trajectories and market volatility.
method Three-stage hierarchical RL framework: Q-teacher, diverse RL agents, and minute-level router.
result Significantly outperforms 6 state-of-the-art baselines in profitability.
A new algorithm computes elastic shape distances between curves efficiently.
problem Computing elastic shape distances between curves in high dimensions.
method Dynamic Programming for optimal diffeomorphisms and Kabsch-Umeyama algorithm for optimal rotation matrices.
result Efficient computation of elastic shape distances with improved efficiency for closed curves.
AI predicts stock winners with 2.43 Sharpe ratio, but returns are highly concentrated.
problem Predicting stock returns with AI, focusing on identifying top winners.
method Deployed a state-of-the-art LLM to autonomously search the web for stock attractiveness, avoiding look-ahead bias.
result AI can generate alpha by identifying top winners, but returns are highly concentrated.
Improves diversity of text-to-image models without sacrificing FID.
problem Lack of diversity and tendency to recreate training set images.
method Adds sparse repellency terms to diffusion SDE to guide trajectories away from a reference set.
result Improves diversity of diffusion models with minimal impact on FID.