When the Right Answer Is Missing: An Arithmetic-Dependent Rejection Bottleneck in Jev
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
NumericJev introduces a training‑free numerical decoding algorithm that allows large language models with Jev‑like structured‑choice interfaces to output precise numerical values. The method refines a numerical range using a multiway decision tree, achieving lower mean absolute error than direct selection from a candidate list. Experiments on an arithmetic benchmark show a 2.93 percentage‑point improvement, and a historical‑index study reports a 4.58% mean relative recall error with zero readout error when the value is supplied.
arXiv:2606. 03606v1 Announce Type: cross Abstract: Large language models achieve strong performance on arithmetic reasoning benchmarks, and one common response to arithmetic brittleness is to delegate computation to code.
Large language models achieve strong performance on arithmetic reasoning benchmarks, and one common response to arithmetic brittleness is to delegate computation to code. Yet models are still often used in settings where they must reason directly from natural language, and trustworthy models should solve small-number arithmetic word problems without external tools.
arXiv:2607. 27031v1 Announce Type: new Abstract: Reports on how sparsification, compression, and lottery tickets change model behavior have been mixed in the prior literature, with beneficial effects observed in some studies and adverse effects in others.
arXiv:2609.37647v1 Announce Type: cross Abstract: Jev is a commercial System One model from TypeSafe AI that does not generate text: given a state and typed questions, it returns a choice from fixed...
arXiv:2608. 15565v1 Announce Type: new Abstract: Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking against known answers, which real ticket streams do not provide.