Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

23456890 · Jun 202019922001200920182026
48 results for face synthesis

Unified model for age-invariant face recognition with photorealistic face synthesis.

problem Reliable face recognition across ages remains challenging due to significant intra-class variations.
method Unified deep architecture for cross-age face synthesis and recognition, continuous face rejuvenation/aging, disentangled age-invariant face representations.
result Superior performance on CAFR and other cross-age datasets, promising generalizability to unconstrained face recognition.

Proposes a new GAN architecture for generating data conditioned on partial information.

problem Generating data conditioned on partial ancillary information.
method Introduces a new Adversarial Network architecture and training strategy.
result The proposed method outperforms standard Conditional GANs in generating data under partial conditioning.

Improved visual speech synthesis using adapted ASR acoustic models.

problem Lack of synchronized audio, video, and depth data for speaker-independent speech-driven visual speech synthesis.
method Adapted an ASR acoustic model trained on audio-only data to the visual speech synthesis domain.
result Viewers significantly prefer animations generated from the adapted ASR acoustic model.

Synthesizes faces from facial features, invariant to pose and expression.

problem Creating realistic face images from facial features.
method Learning facial landmarks and textures from facial-recognition features, training on frontal, neutral-expression images.
result Generated images are invariant to lighting, pose, and expression.

Polynomial fusion layer improves speech-driven facial animation.

problem Recent facial synthesis relies on low-dimensional representations and concatenation, ignoring higher-order interactions.
method Proposes a polynomial fusion layer to model higher-order interactions of facial encodings.
result Demonstrates improved video quality, audiovisual synchronisation, and blink generation.

MO-PaDGAN generates diverse, high-performance designs with multiple metrics.

problem Challenges in generating diverse, high-performance designs with multiple metrics.
method MO-PaDGAN uses a new Determinantal Point Processes based loss function for probabilistic modeling of diversity and performances.
result MO-PaDGAN expands the design space towards high-performance regions and generates new designs with high diversity and performances.

StackGAN++ generates high-quality images from text descriptions.

problem Generating high-quality photo-realistic images from text descriptions.
method Two-stage and multi-stage generative adversarial networks (GANs) with stacked architecture.
result StackGAN++ significantly outperforms other methods in generating photo-realistic images.

Generative model generates synthetic medical images for data augmentation and anonymization.

problem Imbalanced medical imaging data sets, especially for rare pathologies.
method Generative adversarial network (GAN) trained on two public brain MRI datasets.
result Synthetic images improve tumor segmentation performance and serve as an anonymization tool.

FaceSigns embeds a secret watermark in images to authenticate and detect deepfakes.

problem Realistic image and video manipulation threats, especially deepfakes.
method Semi-fragile watermarking using neural networks, robust to face-swapping but fragile to deepfake manipulations.
result FaceSigns can reliably detect deepfake content with high accuracy.

Proposes a network to predict structured uncertainty distributions for images.

problem Previous methods only predicted diagonal covariance matrices, limiting reconstruction accuracy.
method Learns to predict full Gaussian covariance matrices for efficient sampling and likelihood evaluation.
result Accurately reconstructs ground truth correlated residual distributions and generates plausible high frequency samples.

Paper introduces a method for generating interlocutor-aware facial gestures in dyadic settings.

problem Generating appropriate non-verbal behavior for conversational agents in dyadic settings.
method Probabilistic method using multi-modal cues from the interlocutor to synthesize facial gestures.
result The model successfully leverages multi-modal input from the interlocutor to generate more appropriate behavior.

Neural model synthesizes music with flexible timbre controls.

problem Creating audio samples with varied timbres from musical scores.
method Recurrent neural network conditioned on learned instrument embedding followed by WaveNet vocoder.
result Learned embedding space captures diverse timbres and enables interpolation for morphing.

System uses machine learning and automated reasoning to speed up PBE synthesis.

problem Slow synthesis in PBE due to domain-specific knowledge and large training datasets.
method Preprocess SyGuS PBE problems with a neural network to reduce search space, then use automated reasoning for faster solution.
result System outperforms all competing tools in the 2019 SyGuS Competition for the PBE Strings track by 47.65%.

WaveCycleGAN2 improves speech synthesis quality by reducing aliasing.

problem Human ear can still distinguish synthesized speech from natural speech.
method WaveCycleGAN2 uses generators without down/up-sampling modules and combines discriminators from waveform and acoustic parameter domains.
result WaveCycleGAN2 achieves high-quality speech synthesis with comparable mean opinion scores to natural speech.

Automated synthesis planning from scientific literature using AI.

problem Accelerate materials design and discovery by connecting scientific literature to synthesis insights.
method Word embeddings from language models, named entity recognition, conditional variational autoencoder.
result The model predicts precursors for perovskite materials using historical data.

Deep reinforcement learning optimizes retrosynthetic planning for chemical synthesis.

problem Optimizing chemical synthesis plans from molecular targets to simpler starting materials.
method Deep reinforcement learning to estimate synthesis costs and values of molecules.
result Trained neural networks outperform heuristic approaches in synthesizing unfamiliar molecules.

Framework synthesizes geological images minimizing patch distribution discrepancy.

problem Synthesizing realistic geological images from a single exemplar.
method Uses kernel discrepancies and generative neural networks for efficient synthesis.
result Synthesized images match visual patterns and spatial statistics of the exemplar.

Enhanced Tacotron for Japanese speech synthesis improves naturalness.

problem Challenges in end-to-end Japanese speech synthesis due to pitch accents.
method Extended Tacotron with self-attention to capture pitch accent dependencies.
result Proposed systems show improvements but still lag behind traditional pipeline methods.

Inverse Drum Machine separates drum mixes using transcription and synthesis.

problem Separating individual drum tracks from mixed recordings.
method Analysis-by-synthesis framework combining deep learning and automatic transcription.
result Separation quality comparable to supervised methods requiring isolated stems.

Language models predict inorganic synthesis conditions and temperatures.

problem Limited data and heuristic approaches constrain inorganic synthesis planning.
method Language models without fine-tuning predict precursor conditions and temperatures.
result Language models achieve high accuracy in predicting synthesis conditions and temperatures.

Proposes GANs using Capsule Networks for faster image synthesis.

problem Faster image synthesis with fewer training samples and epochs.
method Capsule Networks for image synthesis using GAN architectures.
result Learn data manifold faster and synthesize visually accurate images.

CrossBeam learns to search more efficiently in program synthesis.

problem Efficiently searching through vast program spaces.
method Trains a neural model to guide program synthesis, combining previously explored programs.
result CrossBeam explores much smaller portions of the program space compared to state-of-the-art methods.

MORL uses program synthesis to improve reinforcement learning policies.

problem Difficult to interpret and impose constraints on learned policies from black-box neural networks.
method Iterative framework combining program synthesis and behavior cloning.
result Programmatic representation allows for high-level modifications leading to improved learning.

AutoDiff combines auto-encoder and diffusion model for realistic tabular data synthesis.

problem Generating realistic synthetic tabular data with heterogeneous features.
method Employing auto-encoder architecture to handle tabular data's complexity.
result Synthetic tables from AutoDiff show good statistical fidelity and perform well in machine learning tasks.

New method evaluates text-to-image synthesis for realism, variety, and semantic accuracy.

problem Lack of metrics revealing semantic accuracy in text-to-image synthesis.
method Uses Inception network representations and t-SNE visualization for semantic evaluation.
result Classification accuracy of generated images to real images' visual concepts correlates with semantic accuracy.