New model resolves signal ambiguities in ill-posed systems.
problem Signal retrieval from indirect measurements with known models.
method Variational generative model that captures signal distribution.
result Retrieves consistent signals with high fidelity.
New linear spectral estimators improve phase retrieval accuracy.
problem Recovering vectors from magnitude measurements.
method Linear Spectral Estimators (LSPEs) for phase retrieval.
result LSPEs provide accurate initialization vectors and sharp error bounds.
Two-stage risk control for ranked retrieval systems.
problem Assessing prediction uncertainty and risk control in sequential machine learning systems.
method Developed two-stage risk control methods based on LTT and CRC frameworks, leveraging sequential nature of retrieval and ranking phases.
result The proposed methods provide theoretical guarantees and reduce computational burden compared to prior work.
System optimizes retrieval for personal assistants using reinforcement learning.
problem Needing a neural retrieval-based Q&A system for user memory.
method Direct optimization of F1-score using reinforcement learning.
result Improved retrieval performance on test sets.
Deep learning features improve CBIR system performance.
problem Retrieving similar images from a large database.
method Using features from pre-trained deep learning models for similarity retrieval.
result Significantly superior retrieval results compared to traditional methods.
System converts 3D lung nodule images into embeddings for retrieval.
problem Retrieving similar 3D lung nodule images for radiologist decision support.
method 3D deep learning, semantic representation, transfer learning, similarity score.
result System can measure similarity between nodule annotations and CBIR results.
Paper optimizes AUC for better information ranking.
problem Evaluating retrieval system performance using AUC.
method Non-linear approach using additive regression trees, focusing on multi-class AUC.
result Non-linear approach performs better on multi-relevance datasets.
DINOSAUR improves retrieval by accounting for embedding uncertainty in recommender systems.
problem Retrieval bias towards popular items due to noisy embeddings.
method Samples multiple embeddings per item and queries with sampled embeddings to account for uncertainty.
result Improves coverage of long-tail niche content without sacrificing recall.
New cooperative dynamics enhances retrieval performance in neural networks.
problem Understanding emergent computational capabilities in disordered systems.
method Leveraging statistical mechanics, extended neural network architecture for hetero-associative memory.
result Layers trained with less informative datasets develop retrieval regions of the same amplitude, leading to optimal performance.
Deep Retrieval learns a retrievable structure for efficient large-scale recommendations.
problem Efficiently retrieving top relevant candidates in large-scale recommendation systems.
method Deep Retrieval learns a retrievable structure directly from user-item interaction data, encoding candidates into a discrete latent space and optimizing a model to maximize accuracy.
result Deep Retrieval achieves almost the same accuracy as brute-force baseline and significantly outperforms ANN baselines in a live production system.
prDeep uses a deep neural network to robustly retrieve phases from noisy data.
problem Noise limits traditional phase retrieval algorithms' performance.
method Regularization-by-denoising framework and convolutional neural network.
result prDeep is robust to noise and can handle various system models.
Unified framework for scalable optimization of ranking-based objectives.
problem Scalability issues in optimizing ranking-based performance metrics.
method Unified framework using building block bounds for scalable optimization.
result Substantial improvement in performance over accuracy-objective baseline.
HybridRAG combines KGs and vector retrieval for financial document Q&A.
problem Challenges in extracting and interpreting financial text data.
method Integrates Knowledge Graphs and Vector Retrieval Augmented Generation.
result HybridRAG outperforms traditional methods in Q&A systems for financial documents.
Improves search performance by transferring knowledge from recommender system.
problem Cold start and feedback loop problems in search retrieval.
method Zero-Shot Heterogeneous Transfer Learning framework.
result Significant improvements in relevance and user interactions over production system.
Paper introduces RiskEmbed, a finetuned model for financial risk management.
problem Improving retrieval accuracy in financial question-answering systems.
method Curated dataset and finetuned BERT model for financial domain.
result RiskEmbed significantly outperforms general-purpose and financial embedding models.
Most content-based image retrieval systems consider either one single query, or multiple queries that include the same object or represent the same semantic information. In this paper we consider the content-based image retrieval problem for multiple query images corresponding to different image semantics. We propose a…
Generative memory model avoids vanishing gradients to robustly retrieve patterns.
problem Robust retrieval of stored patterns in the presence of interference and noise.
method Training a generative distributed memory without explicitly simulating attractor dynamics, using a likelihood-based Lyapunov function.
result The model converges to correct patterns upon iterative retrieval and achieves competitive performance as a memory model and a generative model.
The paper introduces subgraph nomination for finding similar subgraphs in networks.
problem Finding similar subgraphs in networks using example subgraphs.
method Formalizes subgraph nomination framework with user-supervised retrieval.
result User-supervised retrieval improves performance in subgraph nomination.
Study improves image-caption retrieval by quantifying feature and posterior uncertainty.
problem Improving reliability in image-caption retrieval tasks with deep learning models.
method Quantified feature and posterior uncertainty for model averaging and reliability measure in image-caption retrieval.
result Consistent improvement in retrieval performance with different datasets and architectures.
The study examines the retrieval capabilities of RBMs and generalized Hopfield networks under various prior distributions.
problem Characterizing the state of RBMs and Hopfield networks under different prior distributions.
method Equivalence between RBMs and generalized Hopfield networks, analysis of phase transitions, and study of retrieval capabilities.
result The retrieval phase is robust and exists at low load for every pattern distribution.
FAKTA automates fact checking across media sources.
problem Automating fact checking across diverse media sources.
method Unified framework integrating document retrieval, stance detection, evidence extraction, and linguistic analysis.
result FAKTA predicts factuality and provides evidence for claims.
Improved text embeddings enhance retrieval from a knowledge base.
problem Efficiently retrieving relevant paragraphs from a large knowledge base.
method Used Stanford Question Answering Dataset (SQuAD) for open-domain question answering. Compared various text-embedding methods and trained deep residual neural models for retrieval.
result Training deep residual neural models for retrieval purposes significantly improves paragraph recall.
This paper improves search efficiency by augmenting autoencoder encodings.
problem Improving search speed and relevance in information retrieval.
method Gradient Augmented Information Retrieval with Autoencoders and Semantic Hashing.
result Gradient Augmented Search (GSA) enhances TF-IDF-based systems.
Test for fairness in IR systems based on protected variables.
problem Unfairness in IR systems due to correlation with protected variables.
method Statistical test for 'distribution parity' in top-K IR results.
result Ensures fairness in IR systems for all users.
Optimizes tree models for better beam search performance.
problem Beam search causes retrieval performance deterioration in tree models.
method Develops Bayes optimality and calibration under beam search, proposes a novel algorithm for optimal tree model learning.
result Eliminates the training-testing discrepancy in tree models.
A new system combines vision and language for person re-identification.
problem Real-world surveillance lacks visual data for person re-identification.
method Two-stream CNN framework with shared logits, CCA for modalities, multi-modal testing protocol.
result 22% improvement in re-identification performance with multi-modal queries.
New DNN boosts image retrieval efficiency.
problem Efficient compression of visual descriptors for large-scale image retrieval.
method Unsupervised multi-codebook quantization in a DNN architecture.
result Significantly outperforms existing methods on visual descriptor datasets.
EENMF improves e-commerce sponsored search efficiency and effectiveness.
problem Improving efficiency and effectiveness of e-commerce sponsored search.
method End-to-end neural matching framework (EENMF) for vector-based ad retrieval and neural pre-ranking.
result Significantly outperforms baseline in real e-commerce traffic.
Symmetry in inverse problems leads to multiple solutions, but breaking symmetry helps deep learning.
problem Symmetry in physical systems causes multiple solutions in inverse problems, hindering deep learning.
method Careful symmetry breaking on training data helps solve inverse problems and improve deep learning performance.
result Symmetry breaking on training data significantly improves deep learning performance in inverse problems.
Interactive image retrieval system learns from user feedback and unlabeled data.
problem Efficiently retrieve relevant images with minimal user interaction.
method Combines active learning and graph-based semi-supervised learning (GSSL) to use unlabeled data.
result High F1 scores with few relevance feedback rounds on large datasets.
Gradient descent variants improve phase retrieval accuracy.
problem Phase retrieval problem in high-dimensional spaces.
method Gradient descent, stochastic gradient descent, Langevin algorithm, dynamical mean-field theory.
result Stochastic variants of gradient descent achieve better generalization in phase retrieval.
GENRE retrieves entities autoregressively, improving efficiency and accuracy.
problem Retrieving entities from queries efficiently and accurately.
method Autoregressive generation of entity names, reducing memory footprint and improving context encoding.
result Significantly improved performance on entity disambiguation, linking, and retrieval tasks.
Proposes UICR to improve novelty in recommendation systems without sacrificing relevance.
problem Balancing relevance and novelty in recommendation systems is challenging, especially for long-tail items.
method Introduces uncertainty modeling in the matching stage and multi-task modeling of model and index uncertainty.
result Improves novelty without sacrificing relevance, as shown by experimental results and online A/B tests.
Language models can answer questions without external knowledge.
problem How much knowledge can be stored in language models?
method Fine-tuning pre-trained language models to answer questions without external context.
result Fine-tuned models perform competitively with open-domain systems.
Diversifies reply suggestions for IM systems using M-CVAE.
problem Improving diversity of automated reply suggestions in instant messaging systems.
method Formulated a generative latent variable model with Conditional Variational Auto-Encoder (M-CVAE) to diversify responses.
result Increased diversity by ~30-40% without significant impact on relevance.
LLM generates coherent macroeconomic stress scenarios for portfolio risk assessment.
problem Macro-financial stress testing and portfolio risk assessment using traditional methods.
method Hybrid prompt-RAG pipeline combining structured prompting and retrieval of country fundamentals and news.
result LLM-generated scenarios yield stable tail-risk amplification with limited sensitivity to retrieval choices.
HabitatAgent offers a multi-agent system for transparent housing consultation.
problem Opaque reasoning and brittle multi-constraint handling in housing recommendation systems.
method HabitatAgent is a multi-agent architecture with specialized roles for memory, retrieval, generation, and validation.
result HabitatAgent achieves 95% accuracy in real user consultation scenarios, significantly outperforming a strong baseline.
Qwant Research improves clinical case matching and information retrieval.
problem Matching and retrieving relevant clinical cases and discussions.
method Approach based on language models and preprocessings, information extraction system using neural networks and linguistic analysis.
result Very encouraging results in information extraction accuracy.
New method retrieves most interfered samples for continual learning.
problem Challenges in continual learning with online data streams.
method Controlled sampling of most interfered samples for replay.
result Consistent gains in performance and reduced forgetting.
Study phase transitions in RBMs with generic priors.
problem Understanding phase transitions in RBMs with various priors.
method Complete analysis of phase diagram, focusing on retrieval phase and paramagnetic phase boundary.
result Retrieval robustness for a wide range of priors and optimal training set size for generalization.
Detects anomalies in product health metrics at eBay for better alerts.
problem Detecting anomalies in unsupervised product health metrics at eBay.
method Developed a Moving Metric Detector (MMD) for anomaly detection and a point-wise ranking model for alert retrieval.
result Improves alert precision and avoids alert spamming in eBay production.
New neural networks model complex phenomena with fewer parameters.
problem Challenges in studying higher-order interactions in neural networks.
method Introducing curved neural networks using the maximum entropy principle.
result Curved neural networks accelerate memory retrieval and exhibit explosive phase transitions.
S2M optimizes mining for diverse data subpopulations.
problem Scalability and uniformity in training sets with many labels and diverse data.
method Doubly-stochastic mining (S2M) computes per-example and minibatch losses on hardest labels/examples.
result S2M ensures good performance across all data subpopulations.
Gradient descent with random initialization solves phase retrieval problems efficiently.
problem Solving systems of quadratic equations for phase retrieval.
method Gradient descent with random initialization for nonconvex least squares problem.
result Gradient descent achieves near-optimal computational and sample complexities for phase retrieval.
Active learning selects high-quality examples for text-to-SQL systems.
problem Efficiently annotate large language models for text-to-SQL systems.
method Formalizes example selection as a constrained experimental design problem over semantic query embeddings, proposing a stratified greedy algorithm that maximizes heteroscedastic mutual information.
result Proposed method significantly reduces labeling effort while maintaining high text-to-SQL retrieval accuracy.
Optimizes ad pruning in sponsored search systems using reinforcement learning.
problem How to efficiently select top K ads from N candidates to maximize revenue.
method Model-free reinforcement learning approach considering downstream as a black-box environment.
result Remarkable improvements in revenue achieved through reinforcement learning.
Proposes a new query autocompletion method that maximizes retrieval performance.
problem Users often select suboptimal queries due to unknown best retrieval performance.
method Formulates query autocompletion as ranking item rankings, uses counterfactual learning.
result Empirical results show improved query suggestions for better retrieval performance.
We introduce MosAIc, an interactive web app that allows users to find pairs of semantically related artworks that span different cultures, media, and millennia. To create this application, we introduce Conditional Image Retrieval (CIR) which combines visual similarity search with user supplied filters or "conditions". …