Rule induction explains neural network predictions globally.
problem Understanding and explaining the behavior of trained models.
method Calculate feature importance, transform inputs, simplify space, fit rule induction model.
result Rule sets explain neural network predictions with 0.80 macro-averaged F-score.
CSTEM models document topics using VAE with semantic distance.
problem Inability of previous topic models to explain semantic relations correctly.
method Continuous semantic topic embedding model using variational autoencoder and Mahalanobis distance.
result Improves topic coherence and semantic relation explanation.
We propose a method for finding alternate features missing in the Lasso optimal solution. In ordinary Lasso problem, one global optimum is obtained and the resulting features are interpreted as task-relevant features. However, this can overlook possibly relevant features not selected by the Lasso. With the proposed met…
Neural NMF discovers hierarchical topics in multilayer data.
problem Detecting latent hierarchical structure in multilayer data.
method Recursive application of nonnegative matrix factorization (NMF) in layers with backpropagation optimization.
result Neural NMF outperforms other hierarchical NMF methods in synthetic and real-world datasets.
LIME explanations can be uncertain, even for accurate models.
problem Uncertainty in LIME explanations undermines trust in machine learning models.
method Demonstrated two sources of uncertainty in LIME: sampling randomness and varying interpretation quality.
result Uncertainty in LIME explanations is present even in high-performing models.
Supervised topic models utilize document's side information for discovering predictive low dimensional representations of documents. Existing models apply the likelihood-based estimation. In this paper, we present a general framework of max-margin supervised topic models for both continuous and categorical response var…
Proposes a new deep topic model using MBN and Lasso.
problem Difficult optimization problem in topic modeling.
method Multilayer bootstrap network (MBN) for dimension reduction, supervised Lasso for topic word discovery.
result Effectiveness demonstrated on 20-newsgroups and TDT2 corpora.
New method uses hypergraphs to improve semi-supervised learning accuracy.
problem Lack of pairwise relationships between samples in network-based learning.
method Un-normalized hypergraph p-Laplacian semi-supervised learning.
result Significantly improved accuracy compared to existing methods.
Training a source model optimally for its own task is suboptimal for downstream transfer.
problem The optimality of a source model for its own task hinders downstream transfer performance.
method Analyzes L2-SP ridge regression, characterizes transfer-optimal source penalty, and identifies alignment-dependent effects.
result Transfer benefits from stronger source regularization when aligned imperfectly, and from weaker regularization when aligned perfectly.
New algorithms speed up EMD computation by four orders of magnitude.
problem High computational complexity of EMD for discrete probability distributions.
method Data-parallel approximation algorithms for linear time complexity.
result Four orders of magnitude faster than existing methods.
Paper defends LSTM-based text classification models from backdoor attacks.
problem Backdoor attacks in LSTM models cause misclassification of spam or malicious speech.
method Backdoor Keyword Identification (BKI) to identify and exclude poisoned samples.
result BKI method effectively mitigates backdoor attacks in various text classification datasets.
We study the problem of learning a latent tree graphical model where samples are available only from a subset of variables. We propose two consistent and computationally efficient algorithms for learning minimal latent trees, that is, trees without any redundant hidden nodes. Unlike many existing methods, the observed …
Tsetlin Machine improves text categorization accuracy and interpretability.
problem High accuracy and interpretability in medical text categorization.
method Representing text as propositional variables, capturing categories with simple formulae, and using the Tsetlin Machine to learn these formulae.
result Tsetlin Machine outperforms other methods on various datasets, delivering best recall and precision.
New attacks exploit transfer learning to misclassify text models.
problem Misclassification attacks against transfer learned text classifiers.
method Novel attack algorithms using unintended features from teacher models.
result Transfer learning increases vulnerability to misclassification attacks.
The 20/60/20 rule improves risk management and portfolio optimization in finance.
problem Understanding and managing financial data with heavy tails.
method Application of the 20/60/20 rule to stock market data, development of new measures for tail heaviness, and integration into portfolio optimization.
result The 20/60/20 rule enhances robustness and performance in portfolio optimization.
For any positive integer r, we exhibit a knot Kr with (20 × 2 r--1 + 1) crossings whose Jones polynomial V (Kr) is equal to 1 mod-ulo 2 r. Our construction rests on a certain 20-crossing tangle T 20 which is undetectable by the Kauffman bracket polynomial pair mod 2.
Pareto's 80/20 rule follows a Gaussian distribution with twice the mean standard deviation.
problem Understanding variations in the 80/20 rule across different contexts.
method Identifying the statistical distribution of the 80/20 rule and its variations.
result The 80/20 rule follows a Gaussian distribution with a standard deviation twice the mean.
Let J1 be the real form of a complex simple Jordan algebra such that the automorphism group is F4(−20). By using some orbit types of F4(−20) on J1, for F4(−20), explicitly, we give the Iwasawa decomposition, the Oshima--Sekiguchi's Kε−Iwasawa decomp…
Proves Montesinos-Nakanishi 3-move conjecture for links up to 20 crossings.
problem Every link is 3-move equivalent to a trivial link.
method Computational methods, including new code in Regina.
result Proves conjecture for links with up to 20 crossings.
Deep Learning identifies 20 critical proteins linked to FLT3-ITD mutation in leukemia.
problem Identifying critical proteins associated with FLT3-ITD mutation in leukemia.
method Hierarchical Deep Learning network using autoencoders for feature extraction.
result Deep Learning accurately correlates 20 critical proteins with FLT3-ITD mutation (97% accuracy).
Enhances relational reasoning with multi-layer architecture.
problem Limited relational reasoning with shallow architectures.
method Multi-layer relation network architecture.
result Solved all 20 tasks in bAbI 20 QA dataset.
Neural nets solve braid untangling up to length 20.
problem Untangling braids in knot theory and group theory.
method Feed-forward neural networks in reinforcement learning.
result Trained neural networks to untangle braids in minimal moves.
The divergence theorem in its usual form applies only to suitably smooth vector fields. For vector fields which are merely piecewise smooth, as is natural at a boundary between regions with different physical properties, one must patch together the divergence theorem applied separately in each region. We give an elegan…
Right-angled Artin groups have elements with stable commutator length at least 1/20.
problem Understanding the complexity of elements in right-angled Artin groups.
method Elementary geometric argument based on earlier work of Culler.
result Non-trivial elements have stable commutator length at least 1/20 in right-angled Artin groups without triangles.
In Peña (2007), MCMC sampling is applied to approximately calculate the ratio of essential graphs (EGs) to directed acyclic graphs (DAGs) for up to 20 nodes. In the present paper, we extend that work from 20 to 31 nodes. We also extend that work by computing the approximate ratio of connected EGs to connected DAGs, of …
This paper is not ready for public consumption, as the last step (Figure 20) is incorrect.
For any discrete, torsion-free subgroup Γ of Sp(n,1) (resp.\ F4−20) with no parabolic elements, we prove that H4n−1(Γ;V)=0 (resp.\ Hi(Γ;V)=0 for i=13,14,15) for any Γ--module V. The main technical advance is a new bound on the p--Jacobian of the barycenter map of Besson--Cour…
The early layers of a deep neural net have the fewest parameters, but take up the most computation. In this extended abstract, we propose to only train the hidden layers for a set portion of the training run, freezing them out one-by-one and excluding them from the backward pass. Through experiments on CIFAR, we empiri…
A graph is 2-apex if it is planar after the deletion of at most two vertices. Such graphs are not intrinsically knotted, IK. We investigate the converse, does not IK imply 2-apex? We determine the simplest possible counterexample, a graph on nine vertices and 21 edges that is neither IK nor 2-apex. In the process, we s…
This is lecture notes of a talk I gave at the Morningside Center of Mathematics on June 20, 2006. In this talk, I survey on Poincare and geometrization conjecture.
We prove the following: there are infinitely many finite-covolume (resp. cocompact) Coxeter groups acting on hyperbolic space H^n for every n < 20 (resp. n < 7). When n=7 or 8, they may be taken to be nonarithmetic. Furthermore, for 1 < n < 20, with the possible exceptions n=16 and 17, the number of essentially distinc…
DNM learns efficient representations of input data using deep neural maps.
problem Learning efficient representations of input data.
method Unsupervised representation learning and visualization using deep convolutional networks and self-organizing maps.
result DNM can learn efficient representations of the input data, reflecting class characteristics.
The number of closed billiard trajectories in a rational-angled polygon grows quadratically in the length. This paper gives an analogue on K3 surfaces, by considering special Lagrangian tori. The analogue of the angle of a billiard trajectory is a point on a twistor sphere, and the number of directions admitting a spec…
We consider gradient estimates to positive solutions of porous medium equations and fast diffusion equations: ut=Δφ(up) associated with the Witten Laplacian on Riemannian manifolds. Under the assumption that the m-dimensional Bakry-Emery Ricci curvature is bounded from below, we obtain gradient estimates which…
Study estimates risks of nuclear waste storage projects.
problem Cost and schedule risks in nuclear waste storage projects.
method Reference class of 216 past projects for cost risk, 200 for schedule risk.
result Cost and schedule risks are substantial for nuclear waste storage projects.
Germany's tax admin costs likely exceed 20% of total revenue, requiring system improvement.
problem High tax administrative costs in Germany and other jurisdictions.
method Statistical data, surveys, and a novel approach to measure total administrative cost as a percentage of total tax revenue.
result Germany's 2021 tax administrative costs likely exceeded 20% of total tax revenue.
The study classifies complex parallelisable nilmanifolds with unobstructed deformations.
problem Characterizing complex parallelisable nilmanifolds with unobstructed deformations.
method Analyzing Lie algebras associated with nilmanifolds and their verbal ideals.
result There are finitely many complex homotopy types of unobstructed complex parallelisable nilmanifolds up to dimension 19, and infinitely many in dimension 20.
We address the question of the growth of firm size. To this end, we analyze the Compustat data base comprising all publicly-traded United States manufacturing firms within the years 1974-1993. We find that the distribution of firm sizes remains stable for the 20 years we study, i.e., the mean value and standard deviati…
Maximal knotless graphs have at least 74% of their vertices' edges.
problem Characterizing maximal knotless graphs and understanding their edge constraints.
method Analyzing edge maximality and constructing graphs to meet constraints.
result There exists an infinite family of maximal knotless graphs with fewer edges than previously thought.
We study 3-valent maps Mn(p,q) consisting of a ring of n q-gons whose the inner and outer domains are filled by p-gons, for p,q≥3. We describe a domain in the space of parameters p, q, and n, for which such a map may exist. With four infinite sequences of maps - prisms Mp(p≥3,4), $M_4(4,q \g…
The paper classifies natural almost Hermitian structures on Lie groups with minimal conformal leaves.
problem Classifying natural almost Hermitian structures on Lie groups with minimal conformal leaves.
method Analyzing Lie groups with a 2-dimensional conformal foliation and classifying structures based on Lie algebra properties.
result 16 multi-dimensional almost Kähler families, 18 integrable families, and 11 Kähler families were constructed.
Paper introduces Arte-Blue Chip Index for diversifying portfolios with art investments.
problem Evaluating blue-chip art as a viable asset class for diversification.
method Developed Arte-Blue Chip Index tracking top-performing artists over 24 years.
result 20% allocation of blue-chip art in a diversified portfolio increases risk-adjusted returns by 20%.
Extended LSTMs improve volatility prediction by 20%.
problem Predicting asset price volatility with long memory.
method Extended LSTMs with multiple flexible timescales.
result Extended LSTMs outperform rough volatility predictions by 20%.
Mirzakhani studied Riemann surfaces and their spaces.
problem Understanding Riemann surfaces and their spaces.
method Survey of her 20 papers.
result Contributions to understanding Riemann surfaces and their spaces.
DPM-Solver speeds up DPM sampling to 10-20 function evaluations.
problem Slow sampling from Diffusion Probabilistic Models (DPMs).
method Exact formulation of diffusion ODE solutions, using change-of-variable and exponentially weighted integral.
result Generates high-quality samples in 10-20 function evaluations.
Solves C^3 null gluing problem for Einstein vacuum equations.
problem Null gluing of up to third-order derivatives of the metric in Einstein vacuum equations.
method Linear and nonlinear analysis of characteristic data close to Minkowski data.
result Solvable up to a 20-dimensional space of obstructions, 10 of which are novel.
Improved clustering speed for 20 clusters on CIFAR-100 dataset.
problem Training time complexity for VAEs with discrete latent variables is linear in the number of clusters.
method Applied a continuous relaxation to discrete variables in Gaussian Mixture VAE, reducing training time complexity to constant.
result Reduced training time from 47 hours to 6 hours for 20 clusters on CIFAR-100.
Hybrid framework predicts Arctic permafrost decline, risks infrastructure, and provides tools.
problem Tackles permafrost decline and infrastructure risk assessment in Arctic territories.
method Hybrid physics-machine learning framework integrating 2.9 million observations.
result Projects mean permafrost fraction decline of -20.3 pp under RCP8.5 forcing, with high-risk zones identified.