arXiv AI

MatLoom: Layered Text-to-Material Generation in a Compact Program Space

arXiv AI
Sep 1

Imag-Eval: a language-grounded framework for interpretable Text-to-Image instruction following evaluation

Imag‑Eval is a new language‑grounded benchmark for evaluating Text‑to‑Image models, focusing on how well they follow compositional natural‑language instructions. It disentangles prompt length from compositional difficulty by independently varying the number of instances and the combination of constraints (rules), providing 1,140 prompts and 8,842 rule combinations. The study shows that for structured skills, the difficulty is mainly driven by the number of grounded rules and their binding to instances rather than prompt length alone.

By Ibrahim Mohamed Serouis, David Jaramillo Duque
arXiv AI
Aug 24

Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality

The paper investigates how small lexical changes in prompts can cause large performance swings in large language models. Using a dataset of 132,000 prompt variants, the authors uncover a scaling law linking higher average task performance to lower variance and greater robustness. They identify domain-specific terminology and explicit action directives as key linguistic factors that stabilize prompts, and propose an automated Prompt-Refining Agent that reduces performance variance by 40.7% in code generation while maintaining or improving mean performance.

By Qipeng Xie, Zi Liang, Jiafei Wu, Yufei Chen, Weizheng Wang, Wenao Ma, Zhong Ming, Haiqin Yang, Kaishun Wu
arXiv Machine Learning
Aug 27

Narcissus: Program Synthesis Using Context-Aware LLM Approximations

Narcissus is a program synthesizer that uses context‑aware large language model (LLM) approximations to guide enumerative search. Unlike prior methods that convert LLM proposals into rule frequencies and lose structural information, Narcissus retains proposals as syntax trees and scores each program expansion based on its surrounding context, ensuring every rule remains reachable. Across five domains and two search backends, it consistently outperforms static guidance and LLM re‑prompting, solving 40% of ARC tasks that raw proposals only solve 13%, all without any LLM calls during search.

By Tilman Hinnerichs, Sebastijan Dumancic, Neil Yorke-Smith
arXiv Machine Learning
Jun 9

IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation

arXiv:2601. 04498v2 Announce Type: replace Abstract: Infographics are composite visual artifacts that combine data visualizations with textual and illustrative elements to communicate information.

By Yinghao Tang, Xueding Liu, Boyuan Zhang, Tingfeng Lan, Yupeng Xie, Jiale Lao, Yiyao Wang, Haoxuan Li, Tingting Gao, Bo Pan, Luoxuan Weng, Xiuqi Huang, Minfeng Zhu, Yingchaojie Feng, Yuyu Luo, Wei Chen