
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
Get every episode summarized
Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
🤗 Upvotes: 31 | cs.CV
Authors:
Yang Yue, Fangyun Wei, Tianyu He, Jinjing Zhao, Zanlin Ni, Zeyu Liu, Jiayi Guo, Lei Shi, Yue Dong, Li Chen, Ji Li, Gao Huang, Dong Chen
Title:
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
Arxiv:
http://arxiv.org/abs/2605.14333v1
Abstract:
Text and faces are among the most perceptually salient and practically important patterns in visual generation, yet they remain challenging for autoregressive generators built on discrete tokenization. A central bottleneck is the tokenizer: aggressive downsampling and quantization often discard the fine-grained structures needed to preserve readable glyphs and distinctive facial features. We attribute this gap to standard discrete-tokenizer objectives being weakly aligned with text legibility and facial fidelity, as these objectives typically optimize generic reconstruction while compressing diverse content uniformly. To address this, we propose InsightTok, a simple yet effective discrete visual tokenization framework that enhances text and face fidelity through localized, content-aware perceptual losses. With a compact 16k codebook and a 16x downsampling rate, InsightTok significantly outperforms prior tokenizers in text and face reconstruction without compromising general reconstruction quality. These gains consistently transfer to autoregressive image generation in InsightAR, producing images with clearer text and more faithful facial details. Overall, our results highlight the potential of specialized supervision in tokenizer training for advancing discrete image generation.
Get every episode summarized
Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Know when Tianyu He turns up
Follow Tianyu He and once a week we email you every new episode they appeared on — including guest spots the show notes never mention, because we read the transcript.
Follow Tianyu HeFree. Pick your own day and time.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Daily Paper Cast

Raven: The Harness of Harnesses for Composable Agentic Intelligence
Daily Paper Cast

MaLiang-Harness: A Programmable Path to Image and Video Generation
Daily Paper Cast

PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation
Daily Paper Cast

Omni-IO Skills: Harnessing Your Agent Omni-Native
Daily Paper Cast