arXiv AI By Yan Wang, Xinyi Hou, Junjun Si, Yanjie Zhao, Weiguo Lin, Haoyu Wang

LaQual: An Automated Framework for LLM App Quality Evaluation

Read the original on arXiv AI →

arXiv:2508. 18636v2 Announce Type: replace-cross Abstract: Representing a new paradigm in software distribution, LLM app stores are rapidly emerging, offering users diverse choices for content generation, coding assistance, education, and more.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Aug 11

HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection

arXiv:2608. 08069v1 Announce Type: cross Abstract: Foundation models are increasingly reused as software components, making model selection a critical software-engineering decision.

By Alireza Joonbakhsh (Shiraz University), Arda Canser Adal{\i} (Utrecht University), Slinger Jansen (Utrecht University), Farshad Khunjush (Shiraz University), Siamak Farshidi (Wageningen University,Research)
arXiv AI
Jun 26

Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

arXiv:2606. 27226v1 Announce Type: new Abstract: Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug.

By Sangwoo Cho, Kushal Chawla, Pengshan Cai, Zefang Liu, Chenyang Zhu, Shi-Xiong Zhang, Sambit Sahu