arXiv AI By Yan Wang, Xinyi Hou, Junjun Si, Yanjie Zhao, Weiguo Lin, Haoyu Wang

LaQual: An Automated Framework for LLM App Quality Evaluation

Read the original on arXiv AI →

arXiv:2508. 18636v2 Announce Type: replace-cross Abstract: Representing a new paradigm in software distribution, LLM app stores are rapidly emerging, offering users diverse choices for content generation, coding assistance, education, and more.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 23

WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective

arXiv:2609.15387v3 Announce Type: replace-cross Abstract: Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automat...

By Chenxu Liu, Zilu Zou, Peizhong Gao, Jiawen Tao, Zhexin Zhang, Guang Chen, Haowei Lin, Ying Zhou, Tianyi Bai, Dolly Deng, Suncong Zheng, Maxm Pan
arXiv AI
Aug 11

HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection

arXiv:2608. 08069v1 Announce Type: cross Abstract: Foundation models are increasingly reused as software components, making model selection a critical software-engineering decision.

By Alireza Joonbakhsh (Shiraz University), Arda Canser Adal{\i} (Utrecht University), Slinger Jansen (Utrecht University), Farshad Khunjush (Shiraz University), Siamak Farshidi (Wageningen University,Research)