LANTERN: Language Model Assessment on Noisy and Transformed Tasks for Understanding Error and Robustness Nuances
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608. 13430v1 Announce Type: cross Abstract: Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence.
arXiv:2410. 07809v2 Announce Type: replace-cross Abstract: Multilingual instruction tuning (MIT) is challenged by the curse of multilinguality, data scarcity, and high computational cost.
arXiv:2511.22341v2 Announce Type: replace-cross Abstract: Previous works identify sensitivity to option order as a key issue in multiple-choice VQA (MC-VQA) evaluation and propose protocols to mitiga...
arXiv:2601. 02580v2 Announce Type: replace-cross Abstract: Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing to collect student performance data for Item Response Theory (IRT) calibration.
The paper proposes a response‑free method for estimating difficulty of reading‑comprehension multiple‑choice items by fine‑tuning a transformer on item wording. It introduces two extensions to a baseline joint‑encoding model: a component‑wise variant that encodes passage, question, and options separately, and a multi‑task variant that adds a question‑answering auxiliary task. Experiments on a corpus of nearly 30,000 items show that both extensions outperform the baseline, especially the multi‑task variant across all metrics and the component‑wise variant in rank ordering, even with limited training data.
arXiv:2511. 21692v3 Announce Type: replace-cross Abstract: We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curation and evaluation.