arXiv:2608. 14329v1 Announce Type: cross Abstract: Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute.
By Dipankar Sarkar
arXiv:2606. 20255v1 Announce Type: cross Abstract: We introduce the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema for Nigerian public discourse that separates surface sentiment from true communicative intent.
By Celestine Achi
arXiv:2608. 09834v1 Announce Type: cross Abstract: Financial sentiment analysis converts unstructured financial news into quantitative signals that can support market analysis and decision-making.
By Fan Zhang, Jiaming Li
arXiv:2608. 09093v1 Announce Type: cross Abstract: How a document's arrangement is written down, its notation, is a training variable that no dataset card records.
By E. M. Freeburg
arXiv:2606. 07897v2 Announce Type: replace Abstract: Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user.
By Alejandro Botas, Paul de Font-Reaulx, Luke Hewitt
Large language models (LLMs) can synthesize financial narratives but may express high confidence when evidence is sparse, stale, or contradictory. This failure is especially consequential in forecasting, where filings, news, prices, volume, and technical signals can disagree.
arXiv:2606. 07897v1 Announce Type: new Abstract: Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user.
By Alejandro Botas, Paul de Font-Reaulx, Luke Hewitt
arXiv:2607. 14174v1 Announce Type: new Abstract: Financial sentiment extraction has largely relied on news text and supervised extraction against return labels alone, leaving 10-K filings -- and volatility, the target risk disclosure is arguably best suited to informing -- comparatively unexplored.
By Sanggyu Sean Choi
arXiv:2606. 15566v1 Announce Type: cross Abstract: Qualitative coding is central to social science, but expert annotation is difficult to scale.
By Eyup Engin Kucuk, Tarik Kelestemur, \"Omer Da\u{g}lar Tanrikulu
arXiv:2605. 22714v3 Announce Type: replace Abstract: Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing through one conversation.
By Sid-Ali Temkit
arXiv:2608. 04200v1 Announce Type: cross Abstract: Financial sentiment classifiers are commonly evaluated against human labels, but strong linguistic performance does not necessarily imply economically useful return predictability.
By Fusheng Luo
arXiv:2608. 05889v1 Announce Type: cross Abstract: Large language models (LLMs) can leave small stylistic traces in text written with their help.
By Przemys{\l}aw Czuma (Polish Association for Artificial Intelligence in Medicine)