Bayesian Quadrature improves ensembling for neural networks with dispersed likelihood peaks.
problem Ensembling neural networks struggles with dispersed, narrow peaks in likelihood surfaces.
method Uses Bayesian Quadrature to construct weighted ensembles of architectures.
result Empirically outperforms state-of-the-art baselines in test likelihood, accuracy, and expected calibration error.
Improved peak detection in ChIP-seq data reduces over-dispersion.
problem Over-dispersion in ChIP-seq data reduces peak detection accuracy.
method Supervised segmentation models with alternative noise assumptions.
result Improved peak detection accuracy compared to natural assumptions.
We establish several new stylised facts concerning the intra-day seasonalities of stock dynamics. Beyond the well known U-shaped pattern of the volatility, we find that the average correlation between stocks increases throughout the day, leading to a smaller relative dispersion between stocks. Somewhat paradoxically, t…
A machine learning model for PMD compensation in dual-polarization systems.
problem Compensating for polarization-mode dispersion (PMD) in dual-polarization systems.
method Model-based machine learning approach using the split-step Fourier method for the Manakov-PMD equation.
result The model converges to within 1% of peak dB performance after 428 iterations, achieving a 0.30 dB reduction in effective signal-to-noise ratio compared to PMD-free case.
The paper links labor income risk to stock returns using industry portfolio returns.
problem Understanding the impact of sectoral shifts on stock returns.
method Using cross-industry dispersion (CID) as a proxy for unemployment risk, the paper examines the relationship between stock returns and the sensitivity of returns to CID innovations.
result Stocks with high sensitivity to CID have lower expected returns, suggesting they are more exposed to sectoral shifts and unemployment risk.
I report a new statistical distribution formulated to confront the infamous, long-standing, computational/modeling challenge presented by highly skewed and/or leptokurtic ("fat- or heavy-tailed") data. The distribution is straightforward, flexible and effective. Even when working with far fewer data points than are rou…
Skewness dispersion predicts future stock market returns, especially in months with monetary policy announcements.
problem Predicting future stock market returns using skewness dispersion.
method Cross-sectional analysis of firm-level realized skewness and stock market returns.
result Skewness dispersion is a significant predictor of future stock market returns, robust to various estimation methods.
New dispersion indices based on inaccuracy and divergence introduced for information measures.
problem Measuring variability in uncertainty measures.
method Introducing new dispersion indices based on Kerridge inaccuracy and Kullback-Leibler divergence.
result Properties, bounds, and examples of new dispersion indices presented.
New heat dispersion laws established for smooth compact manifolds.
problem Understanding heat dispersion in smooth compact manifolds.
method Established new heat dispersion laws through Theorem 1.1 and explored them further with Propositions 3.1 and 3.2.
result New heat dispersion laws for smooth compact manifolds.
MallowsPO enhances LLM fine-tuning with a dispersion index of human preferences.
problem Lack of diversity in human preferences in DPO.
method Developed a dispersion index based on Mallows' theory to characterize preference diversity.
result Demonstrated improved performance in various tasks using the dispersion index.
Geometric focusing affects dispersive estimates for Schrödinger and wave equations.
problem Long-time decay rate in dispersive estimates for Schrödinger and wave equations on non-trapping asymptotically conic manifolds and exact metric cones.
method Classifying the long-time decay rate in dispersive estimates for the Schrödinger and wave equations on non-trapping asymptotically conic manifolds and exact metric cones in terms of the intensity of geometric focusing.
result Each multiplicity of conjugate points within distance π on Y = ∂X0 leads to a |t|1/2-loss in the long-time decay order and a half-order shift in the regularity index in the dispersive estimate for the Schrödinger equation.
Study on stock market volatility and return dispersion during COVID-19.
problem Impact of COVID-19 on stock market volatility and return dispersion.
method Used Google index to proxy epidemic impact, modeled volatility, and analyzed influencing factors of log-return.
result Volatility significantly affected by epidemic and cross-sectional return dispersion, with positive coefficients.
In the recent years, banks have sold structured products such as worst-of options, Everest and Himalayas, resulting in a short correlation exposure. They have hence become interested in offsetting part of this exposure, namely buying back correlation. Two ways have been proposed for such a strategy : either pure correl…
A robust method for decomposing spectral peaks robust to distortion and interference.
problem Decomposing spectral peaks in the presence of distortion and interference.
method Optimizing a nonparametric approach using pseudo-symmetric functions with nonincreasing behavior.
result Decomposed spectral peaks show pseudo-orthogonal behavior and power preserving equality.
The study examines Hawkes processes and their long-term behavior.
problem Understanding the long-term behavior of Hawkes processes.
method Proving functional limit theorems under various conditions on the dispersion of child events.
result Functional limit theorems hold for Hawkes processes with different levels of child event dispersion.
SPADE improves demand forecasting accuracy by 4.5% for post-promotion periods.
problem Overreacting to peak events in demand forecasting leads to biased forecasts.
method SPADE splits forecasting into two tasks: one for peak events and another for post-peak events, using masked convolution filters and a specialized Peak Attention module.
result Overall PPE improvement of 4.5%, 30% improvement for most affected forecasts after promotions and holidays, and 3.9% improvement in PE accuracy.
Urban dispersal events are processes where an unusually large number of people leave the same area in a short period. Early prediction of dispersal events is important in mitigating congestion and safety risks and making better dispatching decisions for taxi and ride-sharing fleets. Existing work mostly focuses on pred…
We find empirically a characteristic sharp peak-flat trough pattern in a large set of commodity prices. We argue that the sharp peak structure reflects an endogenous inter-market organization, and that peaks may be seen as local ``singularities'' resulting from imitation and herding. These findings impose a novel strin…
New framework controls statistical dispersion for high-stakes applications.
problem Understanding and controlling the dispersion of loss distributions in high-stakes applications.
method Simple yet flexible framework for distribution-free control of statistical dispersion measures.
result Proposed methods control statistical dispersion measures with societal implications.
Machine learning classifies surface wave dispersion curves from ambient noise.
problem Classifying surface wave dispersion curves from ambient noise.
method Convolutional neural network (U-net) with transfer learning and supervised learning.
result Machine classification nearly identical to human-picked phases.
Bayesian model tackles spatial count data issues with flexible non-parametric techniques.
problem Challenges in traditional parametric models for spatial count data with unbalanced distributions and complex dependencies.
method Bayesian semi-parametric spatial dispersed count model combining non-parametric techniques and adapted count models.
result Demonstrates superior performance in managing dispersion and capturing intricate spatial patterns.
Joint peak detection is a central problem when comparing samples in genomic data analysis, but current algorithms for this task are unsupervised and limited to at most 2 sample types. We propose PeakSegJoint, a new constrained maximum likelihood segmentation model for any number of sample types. To select the number of…
Mass spectrometry (MS) is an important technique for chemical profiling which calculates for a sample a high dimensional histogram-like spectrum. A crucial step of MS data processing is the peak picking which selects peaks containing information about molecules with high concentrations which are of interest in an MS in…
Study on-chain peak shaving to reduce Ethereum transaction costs.
problem Reducing transaction costs in blockchain networks, especially during congested periods.
method Analyzing transaction-level data from multiple firms across various industries to understand scheduling responses and cost management strategies.
result Firms' scheduling responses to congestion vary, leading to different fee savings and residual costs.
The paper explains two distinct peaks in generalization error for neural networks and simpler models, each governed by different factors.
problem Understanding the peaks in generalization error for neural networks and simpler models.
method Analysis of random feature models and comparison with numerical experiments involving deep neural networks.
result The peaks at N=P and N=D are distinct and governed by different factors (noise sensitivity vs. initialization noise). The heuristic identification of peaks from noisy complex spectra often leads to misunderstanding of the physical and chemical properties of matter. In this paper, we propose a framework based on Bayesian inference, which enables us to separate multipeak spectra into single peaks statistically and consists of two steps.…
We consider time-domain digital backpropagation with chromatic dispersion filters jointly optimized and quantized using machine-learning techniques. Compared to the baseline implementations, we show improved BER performance and >40% power dissipation reductions in 28-nm CMOS.
Machine learning techniques have recently received significant attention as promising approaches to deal with the optical channel impairments, and in particular, the nonlinear effects. In this work, a machine learning-based classification technique, known as the Parzen window (PW) classifier, is applied to mitigate the…
Study on billiard trajectories with fixed bounces.
problem Counting periodic trajectories with specific bounces.
method Analyzes two-dimensional dispersive billiard systems.
result Asymptotic growth of primitive periodic trajectories.
Paper transforms a complex equation into simpler forms for analysis.
problem Analyzing a fourth-order dispersive flow equation on Kähler manifolds.
method Developed the generalized Hasimoto transformation to simplify the equation.
result Explicit expressions derived for three examples of compact Kähler manifolds.
Bayesian framework integrates spectral deconvolution with expert reasoning for robust peak estimation.
problem Challenges in extracting meaningful peaks from noisy or complex spectra.
method Bayesian spectral deconvolution coupled with a physical-property regression layer.
result Recovery of weak peaks in poly(lactic acid) IR spectra related to degradation rates.
FLOPART solves peak detection by creating accurate train and test set predictions.
problem Correctly detecting peaks in sequential data.
method Dynamic programming changepoint algorithm with zero train label errors.
result FLOPART provides highly accurate predictions on both train and test sets.
A new model relaxes constraints on exponential dispersion models.
problem Tight conditions on cumulant function limit the class of exponential dispersion models.
method Introduces K-LED model with Legendre cumulant function and Bregman divergence guidance.
result The model allows for easier computation of mean parameter and includes various distributions.
Probabilistic modeling is cyclical: we specify a model, infer its posterior, and evaluate its performance. Evaluation drives the cycle, as we revise our model based on how it performs. This requires a metric. Traditionally, predictive accuracy prevails. Yet, predictive accuracy does not tell the whole story. We propose…
We explore a decomposition in which returns on a large class of portfolios relative to the market depend on a smooth non-negative drift and changes in the asset price distribution. This decomposition is obtained using general continuous semimartingale price representations, and is thus consistent with virtually any ass…
The paper shows how the generalization curve can have multiple peaks, influenced by data and learning algorithm biases.
problem Understanding the generalization behavior of linear regression models under varying parameterizations.
method Analyzes generalization loss in linear regression models with varying parameterizations, both under- and over-parameterized.
result The generalization curve can have an arbitrary number of peaks, and their locations can be controlled.
PEAKS selects key training examples incrementally based on prediction error and kernel similarity.
problem Dynamic data selection in deep learning models.
method Prediction Error Anchored by Kernel Similarity (PEAKS) for incremental data selection.
result PEAKS outperforms existing selection strategies and yields better performance returns as training data size grows.
PEAK tests means of multiple data streams with sequential betting.
problem Testing means of multiple data streams with nonparametric methods.
method Sequential, nonparametric testing using a betting scheme.
result PEAK provides up to 85% reduction in samples for stopping.
Reduced-order model improves LES for atmospheric pollutant dispersion.
problem Accurate near-field pollutant concentration tracking in urban areas.
method Combining POD and GPR for non-intrusive reduced-order modeling.
result Component-by-component optimization captures spatial scales in high-order modes.
Data analysis in high-dimensional spaces aims at obtaining a synthetic description of a data set, revealing its main structure and its salient features. We here introduce an approach providing this description in the form of a topography of the data, namely a human-readable chart of the probability density from which t…
Dropout improves regularization in flexible models for rare features.
problem Understanding theoretical properties of dropout in generalized linear models.
method Theoretical analysis and application to adaptive smoothing with B-splines.
result Dropout prefers rare features in mean and dispersion parameters.
The standard deviation and Gini mean difference order based on tail behavior.
problem Ordering between standard deviation and Gini mean difference for real-valued risks.
method Analysis of the mean excess function of the pairwise difference ∣X−X′∣. result Dominance regimes of SD and GMD are determined by tail behavior of the distribution.
LLMs learn peaked distributions slowly due to power-law losses.
problem Slow convergence of loss in training large language models.
method Systematic analysis of toy models and empirical evaluation of LLMs.
result Power-law time scaling with an exponent of 1/3 for learning peaked distributions.
Finite-time queue peaks in stochastic networks have logarithmic scaling after geometric thresholds.
problem Queue peak laws in stochastic networks with geometric thresholds.
method Self-normalization mechanism
result Logarithmic scaling of queue peaks after geometric thresholds.
Study dispersive estimates for Schrödinger and wave equations on a cone with specific metric.
problem Pointwise decay estimates for Schrödinger and wave equations on a product cone.
method Modified Hadamard parametrix on Y with ε>π to prove dispersive estimates. result Threshold of conjugate radius ε>π for pointwise dispersive estimates. Proposes a new portfolio optimization method considering reward, dispersion, and asymmetry.
problem Capturing fat-tails and asymmetry in asset return distributions.
method Market model with tempered stable distribution; extended mean-variance optimization.
result Closed-form solutions for VaR and CVaR; efficient frontier extended to three dimensions.
During a stock market peak the price of a given stock (i) jumps from an initial level p1(i) to a peak level p2(i) before falling back to a bottom level p3(i). The ratios A(i)=p2(i)/p1(i) and B(i)=p3(i)/p1(i) are referred to as the peak- and bottom-amplitude respectively. The paper show…
We discuss a short-time existence theorem of solutions to the initial value problem for a third order dispersive flow for closed curves into a compact almost Hermitian manifold. Our equations geometrically generalize a physical model describing the motion of vortex filament. The classical energy method cannot work for …