Towards Data Science

We Pinned Our Model Version to Stay Safe. The Provider Deprecated It Anyway.

arXiv AI
Sep 18

Trust, but Validate the Instrument: Auditing AI-Generated RTL Verification Plans on Authored Security-Regression Proxies

The paper introduces SecTB-RTL, an auditable framework for evaluating AI-generated RTL verification plans against 31 tasks and 124 hardware‑security regressions. In a confirmatory run, the AI model’s responses were rejected by the provider’s schema, and after a schema‑only repair, only nine of 1,857 accepted responses passed the production semantic validator, revealing a mismatch between generation and execution rules. The study demonstrates that schema acceptance does not guarantee execution validity and provides a benchmark, failure‑preserving contract, incident provenance, and governance controls to prevent misreporting of infrastructure behavior as model behavior.

By Hang Xiao, Chuhong Xu, Kainan Zhou, Gangzhen Qian, Lu Yi
arXiv Machine Learning
Jun 25

LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning

arXiv:2606. 24901v1 Announce Type: new Abstract: Continual learning capability is critical for Industrial LLMs, as deployed models must be continuously updated to meet evolving requirements and environments, rather than repeatedly retrained from scratch.

By Hao Jiang, Enneng Yang, Guojie Zhu, Yibin Chen, Yunkun Xu, Zifu Kou, Jiayi Li, Chong Chen, Zhao Cao, Li Shen
Simon Willison
Aug 23

Anthropic’s best AI model struggles to attract users as cheaper tools thrive

Anthropic’s top AI model is struggling to attract users even as cheaper alternatives thrive. The company’s July revenue is projected at $65 bn, up from $47 bn in May, and it expects Q3 profitability while boasting 6,000 high‑spending customers. In contrast, OpenAI’s revenue has risen 35 % this quarter, spurred by GPT‑5.6, and a Ramp AI index shows Anthropic’s newer models (e.g., Fable) are less popular than older ones like Opus 4.8.