Study explores how dataset breadth and depth affect Siamese Neural Network performance.
problem Impact of dataset breadth and depth on Siamese Neural Network performance.
method Experiments with three keystroke datasets varying breadth and depth factors.
result Increasing dataset breadth improves model performance, while depth's impact varies by dataset type.
Higher CEO career breadth correlates with better firm performance.
problem Limited adaptability in complex environments due to specialization.
method Constructed a Breadth Index from 650 CEOs' cross-domain experience, analyzed using regression.
result Higher Breadth Index CEOs outperform industry peers by 9.8 percentage points.
We characterize language generation with stability and breadth, proving impossibility results.
problem Characterizing and proving impossibility results for language generation with stability and breadth.
method Analysis of existing notions of breadth and stability, proving lower bounds.
result Proven impossibility of generating with higher perplexity or lower hallucination rate for stable generators.
Study curves of constant breadth in a specific 3D manifold.
problem Differential geometry of curves in Walker 3-manifolds.
method Investigate curves of constant breadth using Darboux frame.
result Properties of curves of constant breadth in Walker 3-manifolds.
The benefits of portfolio diversification is a central tenet implicit to modern financial theory and practice. Linked to diversification is the notion of breadth. Breadth is correctly thought of as the number of in- dependent bets available to an investor. Conventionally applications us- ing breadth frequently assume o…
New findings show language models can't simultaneously avoid hallucinations and capture all language richness.
problem Achieving both valid output and full language richness in language generation.
method Investigates language generation within a statistical setting, focusing on consistency and breadth.
result For most collections of candidate languages, a language model cannot simultaneously avoid hallucinations and capture all language richness.
This study compares 6 imitation learning algorithms using a common dataset and hyperparameter budget.
problem Difficulty in comparing different imitation learning algorithms due to varying datasets, base RL algorithms, and evaluation settings.
method Reimplemented and updated 6 different imitation learning algorithms, using a common off-policy algorithm (SAC) and a widely-used dataset (D4RL). Evaluated on a range of expert trajectories.
result GAIL consistently performs well across different sample sizes, while AdRIL performs well with one important hyperparameter to tune and behavioral cloning remains a strong baseline when data is plentiful.
In this paper we analyze, evaluate, and improve the performance of training Random Forest (RF) models on modern CPU architectures. An exact, state-of-the-art binary decision tree building algorithm is used as the basis of this study. Firstly, we investigate the trade-offs between using different tree building algorithm…
We show that the number of unique function mappings in a neural network hypothesis space is inversely proportional to ∏lUl!, where Ul is the number of neurons in the hidden layer l.
Recently two search algorithms, A* and breadth-first branch and bound (BFBnB), were developed based on a simple admissible heuristic for learning Bayesian network structures that optimize a scoring function. The heuristic represents a relaxation of the learning problem such that each variable chooses optimal parents in…
Since the beginning of the 21st century, the size, breadth, and granularity of data in biology and medicine has grown rapidly. In the example of neuroscience, studies with thousands of subjects are becoming more common, which provide extensive phenotyping on the behavioral, neural, and genomic level with hundreds of va…
Extended Thistlethwaite's result on Jones polynomials of quasi-alternating links.
problem Characterizing Jones polynomials of quasi-alternating links.
method Analyzing structure and properties of Jones polynomials for quasi-alternating links.
result Jones polynomials of prime quasi-alternating links have no gaps.
Proposes AEGAN for stable GAN training.
problem Training instability in GANs.
method Four-network model with adversarial and reconstruction losses.
result Stabilizes GAN training and prevents mode-collapse.
New approach to abstract neural network representations using renormalization group.
problem Developing truly abstract representations in neural networks.
method Renormalization group approach to expand representations to encompass broader data sets.
result Representations in neural networks become more abstract as data breadth increases and depth increases.
Simple proof of knot genus theorem using Alexander polynomial.
problem Proving the genus of an alternating knot equals half the breadth of its Alexander polynomial.
method Elementary, self-contained proof using Seifert's algorithm.
result Minimal genus surface obtained from any alternating knot diagram.
Research evaluates data poisoning attacks on regression learning and introduces a new defense strategy.
problem Data poisoning attacks on regression learning threaten model integrity in critical systems.
method Realistic scenarios, novel black-box attack, and evaluation on 26 datasets.
result Mean squared error (MSE) increases to 150% with only 2% poisoned samples.
Deep imagination optimizes decision-making in large trees with limited resources.
problem Optimal planning in large decision trees with limited resources and time.
method Analytical solutions and numerical analysis of sampling capacity allocation.
result Optimal policy is to allocate few samples per level for deep exploration, favoring depth over breadth.
Enumerates knots up to five crossings and describes moves between them.
problem Counting and classifying knots up to a specific number of crossings.
method Generated tables of minimal diagrams and derived moves between knots.
result Conjecture about a lower bound for the triple-crossing number based on Alexander polynomial.
Over the last several years, the use of machine learning (ML) in neuroscience has been rapidly increasing. Here, we review ML's contributions, both realized and potential, across several areas of systems neuroscience. We describe four primary roles of ML within neuroscience: 1) creating solutions to engineering problem…
We give new characterisations of sets of positive reach and show that a closed hypersurface has positive reach if and only if it is of class C1,1. These results are then used to prove new alternating Steiner formulæ for hypersurfaces of positive reach. Furthermore, it will turn out that every hypersurface that sat…
An inverse limit of a sequence of covering spaces over a given space X is not, in general, a covering space over X but is still a lifting space, i.e. a Hurewicz fibration with unique path lifting property. Of particular interest are inverse limits of finite coverings (resp. finite regular coverings), which yield fi…
Develops a new framework for integrating satellite allocations in small portfolios.
problem Feasibility constraints in small portfolios, not return predictability, are the primary concerns.
method A four-layer feasibility framework: physical, economic, structural, and epistemic.
result Closed-form feasibility bounds on satellite size, turnover, and breadth without return forecasts.
WILDS 2.0 expands benchmark datasets for unsupervised adaptation.
problem Leveraging unlabeled data for distribution shifts in real-world applications.
method Curated unlabeled data across various applications, tasks, and modalities.
result State-of-the-art methods perform poorly on WILDS datasets.
We use customer demand data for fashion articles on Myntra, and derive a fashionability or style quotient, which represents customer demand for the stylistic content of a fashion article, decoupled with its commercials (price, offers, etc.). We demonstrate learning for assortment planning in fashion that would aim to k…
Enhances LLMs for predicting stock movements by considering news dissemination and context.
problem Lack of consideration for news dissemination and insufficient contextual data in LLMs for stock price prediction.
method Clusters news for reach assessment, enriches prompts with specific data and instructions, fine-tunes an LLM using the dataset.
result Improves prediction accuracy by 8% compared to existing methods.
This study examines ensembling of diffusion models for improved generative quality.
problem Improving generative quality with ensembling of score-based diffusion models.
method Investigated ensembling of scores from multiple diffusion models on image and tabular data.
result Ensembling scores generally improves model likelihood and score-matching loss but not perceptual quality metrics.
Archimedean copulas are popular in the world of multivariate modelling as a result of their breadth, tractability, and flexibility. A. J. McNeil and J. Nešlehová (2009) showed that the class of Archimedean copulas coincides with the class of multivariate ℓ1-norm symmetric distributions. Building upon their result…
Machine learning methods struggle with geometric data, but shape space analysis provides a framework for studying and analyzing geometric variability.
problem Machine learning methods struggle with geometric data
method Shape space analysis provides a mathematical and computational framework
result Characterizes shape variability, compares geometric objects, and analyzes structural trajectories
Neural execution solves complex graph problems like bipartite matching.
problem Solving complex graph algorithms like maximum bipartite matching.
method Reduces bipartite matching to a flow problem and uses Ford-Fulkerson for maximum flow.
result Neural network achieves optimal matching almost 100% of the time.
A new method uses LLMs to discover causal pathways that affect fairness in machine learning.
problem Discovering fairness-relevant causal pathways in the presence of noise and confounding.
method Hybrid LLM-guided causal discovery framework combining active learning and dynamic scoring.
result LLM-guided methods, including the proposed active, dynamically scored variant, outperform baselines in recovering fairness-relevant structure under noisy conditions.
Improves off-policy evaluation with imperfect annotations.
problem Limited dataset coverage for evaluating new policies.
method Doubly robust estimators combining IS and DM, incorporating counterfactual annotations.
result Using annotations within the DM component yields the most desirable theoretical results.
Deep learning approximates shortest path distances in large graphs.
problem Scaling up shortest path distance computation in large networks.
method Deep learning techniques to approximate distances using vector embeddings.
result Feedforward neural networks with embeddings can approximate distances with low distortion error.
Deep learning improves time series forecasting, outperforming other methods.
problem Improving time series forecasting accuracy.
method Deep learning models for time series prediction.
result Deep learning models consistently outperform other methods in forecasting competitions.
Khovanov homology gaps in quasi-alternating links are shown to be one.
problem Khovanov homology gaps in quasi-alternating links.
method Knight Move Conjecture for quasi-alternating links.
result Length of any gap in Khovanov homology and Jones polynomial of quasi-alternating links is one.
Characterizes adequate links using Jones polynomial and crossing number.
problem Characterizing adequate links.
method Using Jones polynomial and crossing number, proving links are adequate.
result Links with specific polynomial properties are adequate.
Turnover-adjusted IR is always lower than classic IR, suggesting managers can improve performance by limiting turnover.
problem The classic relationship between IR and its determinants does not account for turnover costs.
method Mathematical derivations and simulations considering volatility of information coefficient and portfolio turnover.
result Turnover-adjusted IR is lower and managers can improve performance by limiting turnover.
We introduce Dynamic Planning Networks (DPN), a novel architecture for deep reinforcement learning, that combines model-based and model-free aspects for online planning. Our architecture learns to dynamically construct plans using a learned state-transition model by selecting and traversing between simulated states and…
Graph Neural Networks (GNNs) are a powerful representational tool for solving problems on graph-structured inputs. In almost all cases so far, however, they have been applied to directly recovering a final solution from raw inputs, without explicit guidance on how to structure their problem-solving. Here, instead, we f…
Convex clustering solves a stable optimization problem for clustering.
problem Clustering with stable and scalable solutions.
method Solving a convex optimization problem with a single tuning parameter.
result The optimization problem has a unique global minimizer stable to inputs.
Deep Neural Networks (DNNs) provide state-of-the-art solutions in several difficult machine perceptual tasks. However, their performance relies on the availability of a large set of labeled training data, which limits the breadth of their applicability. Hence, there is a need for new {\em semi-supervised learning} meth…
The Bayesian framework is a well-studied and successful framework for inductive reasoning, which includes hypothesis testing and confirmation, parameter estimation, sequence prediction, classification, and regression. But standard statistical guidelines for choosing the model class and prior are not always available or…
Improves diffusion model performance and efficiency through classical search.
problem Tackles inference-time control in diffusion models.
method Proposes a framework combining local and global search for efficient navigation.
result Significant gains in performance and efficiency across various domains.
New model calculates logarithmic surface diameter.
problem Calculating diameter of random hyperbolic surfaces.
method Exploration process inspired by graph breadth-first search.
result Diameter is logarithmic in surface genus.
ITCA optimizes label combination for ambiguous outcomes in multi-class classification.
problem Ambiguous outcome labels in real-world datasets hinder accurate multi-class classification.
method Information-theoretic classification accuracy (ITCA) and search strategies (greedy, breadth-first) guide label combination.
result ITCA improves prediction accuracy and identifies ambiguous labels across diverse applications.
A new method learns priors for Bayesian optimisation to improve performance.
problem Bayesian optimisation tasks often assume strong similarity, which is violated in many cases.
method Replace strong similarity assumption with shape similarity, learn priors for hyperparameters.
result PLeBO and prior transfer find good inputs in fewer evaluations.
The adaptive processing of graph data is a long-standing research topic which has been lately consolidated as a theme of major interest in the deep learning community. The snap increase in the amount and breadth of related research has come at the price of little systematization of knowledge and attention to earlier li…
TSFMs improve financial forecasting across diverse tasks with strong transferability.
problem Complex nonlinear relationships, temporal dependencies, and limited data in financial time series forecasting.
method Pretraining on diverse time series corpora followed by task-specific adaptation.
result Tiny Time Mixers (TTM) achieved 25-50% better performance on limited data and 15-30% improvements on longer datasets.
We formulate a general framework for competitive gradient-based learning that encompasses a wide breadth of multi-agent learning algorithms, and analyze the limiting behavior of competitive gradient-based learning algorithms using dynamical systems theory. For both general-sum and potential games, we characterize a non…