Propose, Don't Judge: An Anytime-Valid Referee for LLM Agents That Mine Investment Factors
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2609.27051v1 Announce Type: new Abstract: Language-model agents now run the whole of quantitative factor research: they propose investment factors, backtest them, select the survivors and retir...
arXiv:2607. 17765v1 Announce Type: cross Abstract: We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events.
arXiv:2607. 05904v1 Announce Type: new Abstract: Training a language model against its own reference-free judgments (the premise of self-rewarding, self-play, and LLM-as-a-judge pipelines) assumes a model's verdict on a shown answer tracks correctness.
arXiv:2608.23308v1 Announce Type: cross Abstract: An LLM asked for a trading strategy returns three artifacts at once: a natural-language rationale, an executable implementation, and once run, a trac...
An LLM asked for a trading strategy returns three artifacts at once: a natural-language rationale, an executable implementation, and once run, a track record. Whether these are the same object is rare...
The paper introduces RuVerBench, a benchmark with 2,458 instances for evaluating the reliability of Large Language Models acting as judges (LaaJ) in verifying rubric compliance within agentic scenarios such as deep research and agentic coding. It systematically meta‑evaluates frontier LLMs, revealing that even the most advanced models perform well yet still produce substantial noise. The study also examines how prompt design, batching, and majority voting affect verification accuracy, noting that weaker models are more prompt‑sensitive, batched verification trades accuracy for efficiency, and majority voting offers diminishing returns.