We Pinned Our Model Version to Stay Safe. The Provider Deprecated It Anyway.
Related stories
Trust, but Validate the Instrument: Auditing AI-Generated RTL Verification Plans on Authored Security-Regression Proxies
The paper introduces SecTB-RTL, an auditable framework for evaluating AI-generated RTL verification plans against 31 tasks and 124 hardware‑security regressions. In a confirmatory run, the AI model’s responses were rejected by the provider’s schema, and after a schema‑only repair, only nine of 1,857 accepted responses passed the production semantic validator, revealing a mismatch between generation and execution rules. The study demonstrates that schema acceptance does not guarantee execution validity and provides a benchmark, failure‑preserving contract, incident provenance, and governance controls to prevent misreporting of infrastructure behavior as model behavior.
LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning
arXiv:2606. 24901v1 Announce Type: new Abstract: Continual learning capability is critical for Industrial LLMs, as deployed models must be continuously updated to meet evolving requirements and environments, rather than repeatedly retrained from scratch.
Anthropic’s best AI model struggles to attract users as cheaper tools thrive
Anthropic’s top AI model is struggling to attract users even as cheaper alternatives thrive. The company’s July revenue is projected at $65 bn, up from $47 bn in May, and it expects Q3 profitability while boasting 6,000 high‑spending customers. In contrast, OpenAI’s revenue has risen 35 % this quarter, spurred by GPT‑5.6, and a Ramp AI index shows Anthropic’s newer models (e.g., Fable) are less popular than older ones like Opus 4.8.
No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage
arXiv:2608. 15286v1 Announce Type: cross Abstract: We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database state diffs across repeated runs, with no LLM in the measurement path, demonstrated on EnterpriseOps-Gym.
When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting
arXiv:2606. 16465v1 Announce Type: new Abstract: AI agents can now take irreversible actions in operational systems, but agent-caused losses are still not clearly assigned, priced, or transferred.
HindsightBench: A Black-Box Behavioral Audit Protocol for Parametric Hindsight in Time-Indexed LLM Decision Tasks
arXiv:2607. 18867v1 Announce Type: new Abstract: Large language models leak parametric knowledge of realized outcomes into historical financial decision tasks.
GPT-4 API general availability and deprecation of older models in the Completions API
GPT-3. 5 Turbo, DALL·E and Whisper APIs are also generally available, and we are releasing a deprecation plan for older models of the Completions API, which will retire at the beginning of 2024.
Certifying Model Upgrades with Slice-Wise Non-Regression and Incumbent Fallback
arXiv:2609.13714v1 Announce Type: new Abstract: An updated model can improve an aggregate metric while degrading a slice that matters to a downstream user. We study checkpoint selection subject to no...
Model Distillation in the API
Fine-tune a cost-efficient model with the outputs of a large frontier model–all on the OpenAI platform
Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvency
arXiv:2607. 07207v1 Announce Type: cross Abstract: We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5.
From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents
arXiv:2607. 08028v1 Announce Type: new Abstract: Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context.