arXiv:2608. 03501v1 Announce Type: new Abstract: AI for Research (AI4Research) leverages AI to automate and improve scientific workflows.
By Zejun Liu, Jian Wu, Ru Peng, Yuliang Ji, Dongyuan Li, Renhe Jiang, Yue Zhang
arXiv:2604.18616v2 Announce Type: replace-cross
Abstract: LLM coding agents can generate correct GPU kernels, but their performance still trails expert libraries. Reaching peak throughput requires co...
By Haohui Mai, Xiaoyan Guo, Xiangyun Ding, Daifeng Li, Qiuchu Yu, Chenzhun Guo, Cong Wang, Jiacheng Zhao, Christos Kozyrakis, Binhang Yuan
The paper investigates how scaling a team of small language‑model agents affects performance across different orchestration architectures. By testing eight architectures on five short‑answer benchmarks and an executable‑code benchmark, it finds that team scaling yields large gains on arithmetic word‑problem tasks but only modest improvements on multiple‑choice and code generation tasks, with no single architecture dominating all tasks. The authors explain these patterns using a generate‑transform decomposition that separates coverage and transformation effects, showing that arithmetic tasks benefit from both coverage and critic‑guided transformation, while other tasks are limited by saturation or poor conversion.
By Blaz Bertalanic, Carolina Fortuna
arXiv:2607. 14568v1 Announce Type: cross Abstract: A companion study ran a 35B mixture-of-experts model on a 2011 NVIDIA Tesla C2075 (Fermi, sm_20, 6GB) as a GPU-prefill/CPU-decode hybrid, because the 4-bit model did not fit in device memory (arXiv:2606.
By A. C. Opus, J. Q. Lu
The paper reports on a large‑scale verified search experiment using a 30B language model on a laptop, evaluating three operator packages—schematic notebooks, named obstacles, and behavioural repulsion—in a factorial design across nine construction problems. Results show that the full composition of operators closes the seed‑to‑record gap more effectively than any single component, increases construction‑hash diversity, and that memory plus repulsion consistently avoids collapse. A frontier proposer achieves similar gains in far fewer samples, but the search ultimately stalls near a plateau where the reference family is adopted and optimized only when provided as code.
By Roberto I. Ono Filho
arXiv:2607. 11696v1 Announce Type: new Abstract: Self-refinement often fails to strengthen few-shot inductive reasoning in large language models.
By Huan Zhu
arXiv:2607. 14541v1 Announce Type: new Abstract: Existing GPU kernel generation benchmarks draw problems from synthetic or curated sources that diverge from deployed workloads.
By Lingyun Yang, Yuxiao Wang, Shenghao Liang, Linfeng Yang, Daocheng Ying, Chunbo You, Rui Zhang, Luping Wang, Yinghao Yu, Guodong Yang, Liping Zhang
AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of the research process, prior work has focused primarily on code implementation and execution, overlooking the importance of this stage, and no benchmark exists to evaluate AI's ability to conduct systematic experiment design.
arXiv:2609.21263v1 Announce Type: new
Abstract: Automated macro placement remains a fundamental challenge in VLSI physical design. Despite decades of research, existing approaches predominantly optim...
By Qiufeng Li, Chengxuan Wang, Rongqian Chen, Quan Cheng, Yihui Ren, Chia-Tung Ho, David Z. Pan, Tian Lan, Weidong Cao
The paper investigates why reasoning‑augmented text‑to‑image models like GoT‑R1 sometimes fail on compositional prompts. By separating the explicit textual plan from the decoder, the authors show that the decoder faithfully executes the plan while the planner often writes incorrect spatial relations, especially for phrasing‑dependent cues. Editing or replacing the plan improves image quality without retraining, demonstrating the viability of modular planner‑decoder architectures.
By Ashritha Gonuguntla
arXiv:2607. 06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with task-specialized architectures.
By Kabir Moghe, Peter Chin
arXiv:2605. 14084v2 Announce Type: replace-cross Abstract: Code agents must both reason over long-horizon repository state and obey strict tool-use protocols.
By Mingzhi Zhu, Michele Merler, Raju Pavuluri, Stacy Patterson