Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

1.1%2.2%3.4%4.5% · Oct 200019922001200920182026
48 results for multimodal web

WebGUM learns web navigation from multimodal data, outperforming previous methods.

problem Limited generalization from domain-specific models in web navigation.
method Instruction-following multimodal agent trained on vision-language foundation models.
result Significant improvement in web navigation performance on benchmarks.

Reduces memory costs for developing countries by replacing images with text.

problem High memory costs associated with multimodal webpages in developing countries.
method Canonical Correlation Analysis (CCA) to replace high-cost modality (images) with low-cost modality (text).
result Reduces memory costs by at least 83.35% through eye-tracking experiments.

Unsupervised methods have proven effective for discriminative tasks in a single-modality scenario. In this paper, we present a multimodal framework for learning sparse representations that can capture semantic correlation between modalities. The framework can model relationships at a higher level by forcing the shared …

2015-11-19abs ↗pdf ↗

MUTLA dataset analyzes multimodal data for teaching and learning analytics.

problem Lack of a comprehensive multimodal dataset for teaching and learning analytics.
method Presented a large-scale MUTLA dataset with synchronized multimodal data from SAIL.
result Provides insights for predicting student engagement and improving adaptive learning.

Filtering data with a pre-trained model improves multimodal contrastive learning performance.

problem Improving the quality of internet-scale multimodal datasets.
method Characterized the performance of filtered contrastive learning under a bimodal data generation model.
result Data filtering using a pre-trained model reduces contrastive learning error by a factor of η\sqrt{η} in the large ηη regime.

TAMA uses LMMs to detect and interpret anomalies in time series data with few labels.

problem Challenges in manual feature engineering and extensive labeled training data for TSAD.
method Leverages LMMs to convert time series into visual formats for few-shot in-context learning.
result Consistently outperforms state-of-the-art methods in TSAD tasks.

We find an invariant characterization of planar webs of maximum rank. For 4-webs, we prove that a planar 4-web is of maximum rank three if and only if it is linearizable and its curvature vanishes. This result leads to the direct web-theoretical proof of the Poincaré's theorem: a planar 4-web of maximum rank is lineari…

2006-05-04abs ↗pdf ↗

Classifies hexagonal circular 3-webs with cubic polar curves.

problem Classifying hexagonal circular 3-webs with algebraic polar curves of degree three.
method Analyzes hexagonal circular 3-webs on unit sphere with polar points on a twisted cubic.
result Completes the classification of hexagonal circular 3-webs with algebraic polar curves of degree three.

We construct flat 3-webs via semi-simple geometric Frobenius manifolds of dimension three and give geometric interpretation of the Chern connection of the web. These webs turned out to be biholomorphic to the characteristic webs on the solutions of the corresponding associativity equation. We show that such webs are he…

2011-08-09abs ↗pdf ↗

We give various results and applications using the connection (E,)(E,\nabla) associated with a dd-web. Precisely, we exhibit fundamental invariants of the web related to the differential equation of first order which presents the web. They cast some new lights on the connection and its construction, both conceptually an…

2007-02-12abs ↗pdf ↗

We investigate the linearizability problem for different classes of 4-webs in the plane. In particular, we apply a recently found in [AGL] the linearizability conditions for 4-webs in the plane to confirm that a 4-web MW (Mayrhofer's web) with equal curvature forms of its 3-subwebs and a nonconstant basic invariant is …

2002-09-22abs ↗pdf ↗

In the present paper we study geometric structures associated with webs of hypersurfaces. We prove that with any geodesic (n+2)-web on an n-dimensional manifold there is naturally associated a unique projective structure and, provided that one of web foliations is pointed, there is also associated a unique affine struc…

2008-12-11abs ↗pdf ↗

Projective invariants described for 3-web geometry, resolving Gronwall's conjecture.

problem Projective invariants of linear 3-webs and resolving Gronwall's conjecture.
method Projective torsion-free Cartan connections, web leaves as geodesics, algorithm for resolving Gronwall's conjecture.
result Resolved Gronwall's conjecture for specific 3-webs.

Much information available on the web is copied, reused or rephrased. The phenomenon that multiple web sources pick up certain information is often called trend. A central problem in the context of web data mining is to detect those web sources that are first to publish information which will give rise to a trend. We p…

2012-06-27abs ↗pdf ↗

In this paper we study the linearizability problem for 3-webs on a 2-dimensional manifold. With an explicit computation based on the theory developed in the paper "On the linearizability of 3-webs" (Nonlinear analysis 47, (2001) pp. 2643-2654), we examine a 3-web whose linearizability was claimed in the same paper. We …

2006-02-23abs ↗pdf ↗

Veronese webs appear as the natural way of passing to the quotient of curves in the projective space. In thi paper, we give the link between classical multidimensionnal webs and veronse webs by mean of interpolation.

2004-08-18abs ↗pdf ↗

TCT learns multimodal sequence representations by translating from related sequences.

problem Challenges in learning semantic representations from multimodalities.
method Transformer based Cross-modal Translator (TCT) combined with Multimodal Transformer Network (MTN).
result Proposed method achieves new state-of-the-art performance on video-grounded dialogue.

In this paper, we reconstruct Kuperberg's G2G_2 web space. We introduce a new web (a trivalent diagram) and new relations between Kuperberg's web diagrams and the new diagram. Using the G2G_2 webs, we define crossing formulas corresponding to R-matrices associated to some G2G_2 irreducible representations and calculate…

2015-03-29abs ↗pdf ↗

Paper proposes RMFN for multimodal language analysis.

problem Modeling interactions between language, visual, and acoustic modalities.
method Recurrent Multistage Fusion Network (RMFN) decomposes fusion into stages focusing on subsets of multimodal signals.
result RMFN achieves state-of-the-art performance across multimodal sentiment analysis, emotion recognition, and speaker traits recognition datasets.

Gronwall conjecture states that a planar 3-web which admits more than one distinct linearization is locally equivalent to an algebraic web. We give a partial answer to the conjecture in the affirmative for the class of planar 3-webs with the web curvature that vanishes to order three at a point. The differential relati…

2011-02-01abs ↗pdf ↗

We find relative differential invariants of orders eight and nine for a planar nonparallelizable 3-web such that their vanishing is necessary and sufficient for a 3-web to be linearizable. This solves the Blaschke conjecture for 3-webs. As a side result, we show that the number of linearizations in the Gronwall conject…

2004-11-21abs ↗pdf ↗

This paper proposes a model to learn multimodal representations robust to missing data.

problem Learning multimodal representations from heterogeneous sources of information.
method Optimizes a joint generative-discriminative objective across multimodal data and labels, factorizing representations into multimodal discriminative and modality-specific generative factors.
result The proposed model achieves state-of-the-art performance on six multimodal datasets and can reconstruct missing modalities without significant performance drop.

FMT model improves multimodal sequential learning across language, vision, and acoustic data.

problem Modeling spatio-temporal dynamics across multiple modalities.
method Factorized Multimodal Transformer (FMT) that models intramodal and intermodal dynamics in a factorized manner.
result FMT outperforms existing models on 3 datasets and 21 labels, setting new state of the art.

The Gronwall conjecture states that a planar 3-web of foliations which admits more than one distinct linearizations is locally equivalent to an algebraic web. We propose an analogue of the Gronwall conjecture for the 3-web of foliations by Legendrian curves in a contact three manifold. The Legendrian Gronwall conjectur…

2012-02-29abs ↗pdf ↗

Develops a contrastive framework for data-efficient multimodal learning.

problem Expensive training of multimodal generative models requiring related multimodal data.
method Contrastive framework for multimodal learning, distinguishing related from unrelated data.
result Data-efficient multimodal learning on challenging datasets for various VAE models.

We study non-flat planar 3-webs with infinitesimal symmetries. Using multi-dimensional Schwarzian derivative we give a criterion for linearization of such webs and present a projective classification thereof. Using this classification we show that the Gronwall conjecture is true for 3-webs admitting infinitesimal symme…

2014-11-04abs ↗pdf ↗

Plug-and-play multimodal controller improves class-conditional image generation.

problem Generating class-conditional images from user-specified labels.
method Introduces a `multimodal controller` to generate multimodal data without additional learning parameters.
result Multimodal controlled generative models produce higher quality class-conditional images and novel modalities.