arXiv:2609.07901v1 Announce Type: new
Abstract: Weight quantization largely determines the economics of serving open-weight LLMs. Its costs are usually assessed with capability benchmarks, on which 4...
By Dachi Kurtskhalia
arXiv:2509. 11208v3 Announce Type: replace-cross Abstract: Transformers used for evidence-grounded binary adjudication (e.
By Leon Chlon, Ahmed Karim, Maggie Chlon, MarcAntonio Awada
QEncodeBench evaluates whether large language models can translate classical constraint problems into verified quantum phase oracles. The benchmark measures the correctness of generated circuits using an adversarial self‑validated verifier that checks full solution‑set equivalence while enforcing resource limits. Results show that models lacking a reasoning mode perform poorly, whereas enabling native reasoning improves accuracy tenfold; semantic errors dominate, and neuro‑symbolic pipelines close most gaps by delegating critical composition to deterministic procedures.
By Xujun Che, Hanhan Wu, Yuchen Yuan, Chenyang Yu
arXiv:2608. 16391v1 Announce Type: cross Abstract: As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem.
By Xiangfan Wu, Zonghao Ying, Huiyu Wu, Xing Zheng, Huangsheng Cheng, Xiaorong Shi, Jing Guo
arXiv:2604. 07650v2 Announce Type: replace Abstract: The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent?
By Chenchen Kuai, Jiwan Jiang, Zihao Zhu, Hao Wang, Keshu Wu, Zihao Li, Yunlong Zhang, Chenxi Liu, Zhengzhong Tu, Zhiwen Fan, Yang Zhou
arXiv:2608.31105v1 Announce Type: new
Abstract: Users of a deployed language model routinely encounter behaviours that testing almost never surfaces, since deployment puts the model through orders of...
By Adrians Skapars, Edoardo Manino