Language models are typically trained on plain text extracted from documents, discarding layout, figures, and typography that carry meaningful information. This paper investigates training foundation models directly on visual renderings of documents instead of extracted text, testing whether "seeing" a page beats "reading" it. Across multiple model backbones and benchmarks, visual pretraining outperforms text-only pretraining on the same source material. Applications include more efficient and capable foundation models for tasks involving richly formatted documents, web pages, scientific papers, and forms - domains where converting to plain text currently discards valuable structural and visual information.
Authors: Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen
Paper: https://arxiv.org/abs/2607.09657v1
Fler avsnitt av Eye on AI Weekly Research Watch
Visa alla avsnitt av Eye on AI Weekly Research WatchEye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
