Improves neural machine translation by learning better source representations with relation networks.
problem Forgetting distant information and disregarding relationship between source words in neural machine translation.
method Introduces relation networks to learn better source representations by associating source words with each other and retaining their relationships.
result Significantly improves translation performance over conventional encoder-decoder models and outperforms approaches involving supervised syntactic knowledge.
The paper tackles multi-source learning by integrating information from various sources using neural variational inference.
problem Learning from multiple sources of information with challenges in representation and inference.
method Formulated a variational autoencoder framework where each encoder is conditioned on a different source, integrating beliefs via divergence measures.
result Demonstrated that conflict detection and redundancy can increase robustness in multi-source inference.
Enhances source domain knowledge with target data for transfer learning.
problem Limited data in target domains and rigid model assumptions in transfer learning.
method Transfer learning through Enhanced Sufficient Representation (TESR).
result TESR enhances source domain knowledge with target data, improving transfer learning performance.
This work examines a semi-blind single-channel source separation problem. Our specific aim is to separate one source whose local structure is approximately known, from another a priori unspecified background source, given only a single linear combination of the two sources. We propose a separation technique based on lo…
Contrastive Code Representation Learning improves code summarization and type inference.
problem Code representations are sensitive to edits, hindering downstream semantic understanding tasks.
method ContraCode: a contrastive pre-training task that learns code functionality.
result Contrastive pre-training improves code summarization and type inference accuracy.
G5 universal GRAPH-BERT learns graph representations across different datasets.
problem Learning graph representations across diverse graph datasets with distinct input and output configurations.
method G5 introduces a pluggable model architecture with input and output components for each graph data source, connected via a unified layer and fusion layer.
result G5 removes obstacles for cross-graph representation learning and transfer, even for sparse data.
Paper develops upper-bounds for target general loss in multiple source DA and DG settings.
problem Complexity and trade-offs in multiple source domain adaptation and domain generalization.
method Defines two types of domain-invariant representations and studies their pros, cons, and trade-offs.
result Developed upper-bounds for target general loss offer insights into domain-invariant representations.
Deep learning detects software vulnerabilities from source code.
problem Automated detection of software vulnerabilities in source code.
method Deep feature representation learning on lexed source code.
result Deep learning can effectively detect software vulnerabilities.
Transfer knowledge from multiple sources to improve matrix completion.
problem Matrix completion with noisy data.
method Aggregating singular subspaces information from multiple sources to solve a two-way PCA problem and transform into a low-dimensional linear regression.
result Guaranteed statistical efficiency in transforming the high-dimensional target matrix completion problem.
This paper develops source traces for faster TD learning.
problem Improving temporal difference learning speed and generalization.
method Introduces source traces as a backward view of successor representations, enabling TD errors to be propagated to potential causal states.
result Demonstrates faster generalization and improved performance of source traces compared to previous methods.
Improves domain adaptation by clustering target representations.
problem Learning invariant and discriminative representations for unlabeled target domains.
method Simultaneously learns tightly clustered target representations and assigns each cluster to a unique class from the source.
result Achieves state-of-the-art performance in balanced, imbalanced, and partial domain adaptation.
New risk decompositions clarify domain adaptation issues.
problem Domain adaptation challenges with different training and test distributions.
method Representation Bayesian Risk Decompositions, hybrid argument.
result Clarifies factors (2) and (3) as reasons for generalization failure.
Transformer model improves source code summarization.
problem Generating readable summaries of source code.
method Transformer model with self-attention mechanism for code representation.
result Transformer model outperforms state-of-the-art techniques.
New method disentangles sources of different timescales in planetary seismic data.
problem Unsupervised source separation of multi-scale seismic data from planetary missions.
method Wavelet scattering spectra for multi-scale clustering and variational autoencoder for source separation.
result Disentangles sources with different timescales in InSight mission seismic data.
New findings on representation learning beyond linear functions.
problem Achieving diversity in representation learning beyond linear prediction functions.
method Analysis of eluder dimension and empirical risks.
result Diversity holds even with nonlinear prediction functions and multiple layers in neural networks.
BNN learns shared features between two data sources for specific tasks.
problem Learning shared features between two data sources for specific tasks.
method BNN uses two CNNs to project data sources into a feature space and learns a common representation for each task.
result BNN achieves state-of-the-art performance on various tasks.
A new method aligns source and target distributions by tuning their weights.
problem Domain adaptation on unlabeled target datasets using labeled source datasets.
method Weighted Joint Distribution Optimal Transport (WJDOT) method that finds alignment between source and target distributions and re-weighting of source distributions.
result Achieves state-of-the-art performance on simulated and real-life datasets.
Transfer learning aims to improve learning in target domain by borrowing knowledge from a related but different source domain. To reduce the distribution shift between source and target domains, recent methods have focused on exploring invariant representations that have similar distributions across domains. However, w…
The paper provides guarantees for learning nonlinear representations from multiple non-identically distributed data sources.
problem Learning from non-identically distributed and dependent data.
method Established statistical guarantees for learning general nonlinear representations from multiple data sources.
result The excess risk of the estimated function decays as a function of the sample complexity and task diversity.
TopoGeoScore selects robust checkpoints using only source-domain representations.
problem Selecting robust checkpoints without target-domain labels or samples.
method Constructs class-conditional mutual k-nearest-neighbour graphs and extracts three interpretable signals.
result Source representations contain measurable global-local-topological evidence of robustness.
Wavesplit separates speech from mixtures using clustering.
problem Permutation problem in speech separation.
method End-to-end system infers speaker representations and estimates signals.
result Robust separation of long recordings, new benchmarks set.
This paper improves few-shot learning by reducing sample complexity using representation learning.
problem Reducing sample complexity for target tasks with limited data.
method Representation learning to pool all source task samples for target task learning.
result Representation learning can achieve substantial sample size reduction, bypassing the $Ω(rac{1}{T})$ barrier.
Paper proposes a method to learn linear regression models using multiple pre-trained models.
problem Learning a linear regression model with limited target data.
method Representation transfer learning method using multiple pre-trained models.
result The method achieves better sample complexity compared to baseline methods.
This study examines representation bias in open-source Qwen models for investment decisions.
problem Representation bias in financial applications of large language models.
method Balanced round-robin prompting over 150 U.S. equities, constrained decoding, token-logit aggregation.
result Firm size and valuation increase model confidence, while risk factors decrease it.
Paper shows deep nets can't guarantee domain adaptation without invariant features.
problem Learning domain-invariant features for successful domain adaptation.
method Proposes a counterexample and a new generalization upper bound.
result Shows conditional shift and characterizes a fundamental tradeoff.
FlowGN tackles graph representation learning by tracing information flow paths.
problem GCNs struggle with over-smoothing and scalability issues.
method FlowGN introduces a 'SourceoSink' mode and 'information flow path' concept. result FlowGN outperforms state-of-the-art GCNs in public datasets.
The paper tackles transfer learning for growing matrix representations, improving estimation accuracy.
problem Structured matrix estimation under growing ambient dimensions and latent representations.
method Proposes a general transfer framework decomposing target parameters into embedded source components, low-rank innovations, and sparse edits. Develops an anchored alternating projection estimator.
result Establishes deterministic error bounds that separate target noise, representation growth, and source estimation error, yielding improved rates.
Paper introduces a method to create robust representations against covariate shifts.
problem Distribution shift between training and testing data in machine learning.
method Introduces a variational objective with two components: discriminative representation and invariant support.
result Optimal representations ensure robustness to covariate shifts, improving performance on DomainBed.
Models learn to represent edits from natural language and code.
problem Learning distributed representations of edits.
method Combining a neural editor with an edit encoder.
result Models can capture the structure and semantics of edits.
Paper explores understanding of neural source code embeddings.
problem Lack of understanding of contents and characteristics of code2vec embeddings.
method Small case study using code2vec embeddings to create binary SVM classifiers and compare performance with handcrafted features.
result Code2vec embeddings perform similarly to handcrafted features and have more evenly distributed information gains.
MAGIC generates image collages from set templates using attention and set representations.
problem Generating image collages from set templates is challenging for classical models.
method Memory Attentive Generation of Image Collages (MAGIC) using Set-Transformer layers and set-pooling.
result MAGIC can generate image collages from set templates in one forward pass.
Paper tackles robust domain adaptation without target domain data.
problem Learning domain invariant representations without target domain data.
method Integrates deep autoencoder and causal structure learning into a unified model.
result CAE learns causal representations using only source domain data.
New layers estimate complex time-frequency masks without phase wrapping issues.
problem Lack of phase estimation in deep learning-based speech enhancement and source separation.
method Proposes magbook, phasebook, and combook layers for complex mask estimation.
result Match state-of-the-art performance on speaker separation datasets.
Paper tackles entity matching over multi-source data, optimizing alignment and mitigating negative transfer.
problem Learning effective entity matching models over multi-source large-scale data with relaxed assumptions.
method Proposes a Relaxed Multi-source Large-scale Entity-matching (RMLE) problem and Incentive Compatible Pareto Alignment (ICPA) method.
result Optimized cross-source alignments and mitigated negative transfer, improving entity matching accuracy.
Novel bounds for deep MDA algorithms improve performance and efficiency.
problem Improving performance of MDA algorithms with few target labels and pseudo labels.
method Information-theoretic tools and novel deep MDA algorithm.
result Algorithm-dependent generalization bounds for MDA.
Unified approach to homological representations of topological groups.
problem Constructing homological representations of topological groups.
method Functorial approach using topological enrichment of Quillen bracket construction.
result Unified construction of homological representations for mapping class groups and motion groups.
New framework tackles multi-source domain adaptation with optimism and consistency.
problem Adjusting mixture distribution weights and ensuring low error on target domain.
method Mildly optimistic objective function and consistency regularization.
result Beats current state of the art in multi-source domain adaptation.
StrEBM learns distinct latent components for better source separation.
problem Blind source separation with identifiable and decoupled latent components.
method Structured latent energy-based model with learnable structural biases.
result The model effectively recovers source components from mixed signals.
IdBench benchmarks semantic representations of identifiers, revealing strengths and weaknesses.
problem Evaluating semantic representations of identifiers in source code.
method Created a benchmark using developer ratings, evaluated natural language and source code embeddings, and compared lexical string distance functions.
result No single technique provides a satisfactory representation of semantic similarities, but ensemble models can improve performance.
New model learns execution of code using GNNs.
problem Stagnation of computer system performance due to Moore's Law.
method Multi-task GNN over low-level code and program state.
result Improved performance on dynamic tasks (26% and 45% over state-of-the-art).
New method learns robust joint representations by translating between modalities.
problem Learning robust joint representations from noisy or missing modalities.
method Cyclic translations between modalities with cycle consistency loss.
result Achieves state-of-the-art results on multimodal sentiment analysis datasets.
Domain adaptation aims at generalizing a high-performance learner on a target domain via utilizing the knowledge distilled from a source domain which has a different but related data distribution. One solution to domain adaptation is to learn domain invariant feature representations while the learned representations sh…
Proposes a novel framework for unsupervised domain adaptation using causal representations.
problem Transferability of deep model representations across domains is limited.
method Integrates causal inference into deep learning pipeline for domain-invariant feature learning.
result Demonstrates superior performance in unsupervised domain adaptation using causal representations.
Training-free source selection for LLM families with shared vocabularies
problem Source selection for LLM families with shared vocabularies
method Fisher alignment at vocabulary scale
result Fisher alignment is a cosine between kernel mean embeddings in the joint activation-error space
New approach to deep learning for domain adaptation.
problem Learning a model on a target domain using a similar source domain.
method Introducing a search framework for correct alignment of high-level representations.
result Conceptual domain adaptation improves deep learning performance.
ICLR 2021 challenge in computational geometry and topology attracted 16 teams.
problem Designing and evaluating computational methods in differential geometry and topology.
method Designing and hosting an open-source competition with repositories Geomstats and Giotto-TDA.
result 16 teams participated in the challenge, showcasing innovative contributions to computational geometry and topology.
This research improves representation learning for new domains with limited new supervision.
problem Learning representations that generalize well to new domains with minimal new data.
method Encourages linearity of factors of variation through learned linear transformations called latent canonicalizers.
result Reduces the number of observations needed to generalize to a similar target domain compared to supervised baselines.
Myia compiler optimizes ML models with efficient AD for array programming.
problem Efficient automatic differentiation for array programming in ML.
method Introduces a new graph-based IR that supports function calls, higher-order functions, and recursion.
result Myia compiler enables efficient AD using source transformation without a tape, supporting higher-order derivatives.