A new algorithm optimizes unknown functions with noisy data and unmatched features.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Partial soft-matching distance improves neural representation comparison by allowing some neurons to remain unmatched.
Arnold introduced invariants , and for generic planar curves. It is known that both and are invariants for generic spherical curves. Applying these invariants to underlying curves of knot diagrams, we can obtain lower bounds for the number of Reidemeister moves for uknotting.…
Convolutional networks outperform fully-connected ones in certain tasks.
We consider a dynamic market model of liquidity where unmatched buy and sell limit orders are stored in order books. The resulting net demand surface constitutes the sole input to the model. We prove that generically there is no arbitrage in the model when the driving noise is a stochastic string. Under the equivalent …
We consider a dynamic market model where buyers and sellers submit limit orders. If at a given moment in time, the buyer is unable to complete his entire order due to the shortage of sell orders at the required limit price, the unmatched part of the order is recorded in the order book. Subsequently these buy unmatched …
We introduce a family of adaptive estimators on graphs, based on penalizing the norm of discrete graph differences. This generalizes the idea of trend filtering [Kim et al. (2009), Tibshirani (2014)], used for univariate nonparametric regression, to graphs. Analogous to the univariate case, graph trend filteri…
We study the problem of semi-supervised singing voice separation, in which the training data contains a set of samples of mixed music (singing and instrumental) and an unmatched set of instrumental music. Our solution employs a single mapping function g, which, applied to a mixed sample, recovers the underlying instrum…
Crowdsourced labeling recovers task types with minimal queries.
TNDE quantifies dynamic gene drivers from single-cell snapshots.
Colored noise improves neural network robustness against adversarial attacks.
A novel federated learning framework resolves structural misalignment in model fusion.
Graphs model human mobility patterns, reducing errors in data matching.
Few-shot models have become a popular topic of research in the past years. They offer the possibility to determine class belongings for unseen examples using just a handful of examples for each class. Such models are trained on a wide range of classes and their respective examples, learning a decision metric in the pro…
Neural networks have demonstrated unmatched performance in a range of classification tasks. Despite numerous efforts of the research community, novelty detection remains one of the significant limitations of neural networks. The ability to identify previously unseen inputs as novel is crucial for our understanding of t…
Deep ensembles have been empirically shown to be a promising approach for improving accuracy, uncertainty and out-of-distribution robustness of deep learning models. While deep ensembles were theoretically motivated by the bootstrap, non-bootstrap ensembles trained with just random initialization also perform well in p…
Enhances community detection in correlated networks with node attributes.
Vacant taxi drivers' passenger seeking process in a road network generates additional vehicle miles traveled, adding congestion and pollution into the road network and the environment. This paper aims to employ a Markov Decision Process (MDP) to model idle e-hailing drivers' optimal sequential decisions in passenger-se…
Two algorithms tackle heavy-tailed rewards in reinforcement learning with linear function approximation.
Big data sets must be carefully partitioned into statistically similar data subsets that can be used as representative samples for big data analysis tasks. In this paper, we propose the random sample partition (RSP) data model to represent a big data set as a set of non-overlapping data subsets, called RSP data blocks,…
Prevents sensitive data generation in diffusion models using labeled and unlabeled data.
Study reveals Data Shapley's inconsistent performance in data selection tasks.
PRRO generates synthetic tabular data that improves SL performance and class distribution.
Defines data science as a natural ecosystem with challenges and missions.
Differences in data size per class, also known as imbalanced data distribution, have become a common problem affecting data quality. Big Data scenarios pose a new challenge to traditional imbalanced classification algorithms, since they are not prepared to work with such amount of data. Split data strategies and lack o…
Synthetic data enhances analytics but requires careful volume management.
New test ensures quality of shared data in machine learning.
Differential privacy allows quantifying privacy loss resulting from accessing sensitive personal data. Repeated accesses to underlying data incur increasing loss. Releasing data as privacy-preserving synthetic data would avoid this limitation, but would leave open the problem of designing what kind of synthetic data. W…
Paper creates fair synthetic data ensuring equal predictions across sensitive attributes.
DPA preserves data distribution in reduced dimensions.
Efficient synthetic data generation improves model performance on tabular data.
For most problems in science and engineering we can obtain data sets that describe the observed system from various perspectives and record the behavior of its individual components. Heterogeneous data sets can be collectively mined by data fusion. Fusion can focus on a specific target relation and exploit directly ass…
DAERNN models censored data using neural networks with data augmentation.
Data preprocessing techniques are devoted to correct or alleviate errors in data. Discretization and feature selection are two of the most extended data preprocessing techniques. Although we can find many proposals for static Big Data preprocessing, there is little research devoted to the continuous Big Data problem. A…
Data stream classification methods demonstrate promising performance on a single data stream by exploring the cohesion in the data stream. However, multiple data streams that involve several correlated data streams are common in many practical scenarios, which can be viewed as multi-task data streams. Instead of handli…
Data collection is a major bottleneck in machine learning and an active research topic in multiple communities. There are largely two reasons data collection has recently become a critical issue. First, as machine learning is becoming more widely-used, we are seeing new applications that do not necessarily have enough …
This paper quantifies uncertainty in Data Shapley using statistical inference.
DCoM uses deep neural networks to detect semantic data types from raw column values.
Causal inference from observational data is the goal of many data analyses in the health and social sciences. However, academic statistics has often frowned upon data analyses with a causal objective. The introduction of the term "data science" provides a historic opportunity to redefine data analysis in such a way tha…
Task-agnostic data valuation without validation requirements.
Data mining is about obtaining new knowledge from existing datasets. However, the data in the existing datasets can be scattered, noisy, and even incomplete. Although lots of effort is spent on developing or fine-tuning data mining models to make them more robust to the noise of the input data, their qualities still st…
New algorithm improves data imputation for complex multimodal data sets.
New algorithms for clustering and synthetic data generation of heterogeneous tabular datasets.
WeMix improves data augmentation by correcting bias in deep learning.
A new method reduces data valuation variance for more trustworthy data trading.
Adapts data analysis for growing data, improving generalization guarantees.
This work redefines data-centric AI by unifying categorical and cochain notions.
SMOTE-DP enhances synthetic data privacy without sacrificing utility.