← Back to all news
Simon Willison August 16, 2026

Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things

Read the original on Simon Willison →

Friday's big release was Qwen 3. 8 27B , an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba's Qwen research lab.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Simon Willison.

  • llms
  • agents
  • efficiency
  • benchmarks
  • safety
  • data-engineering

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jun 11

Energy-Efficient On-Device RAG on a Mobile NPU: System Design and Benchmark on Snapdragon X Elite

arXiv:2606. 11257v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) pipelines are compute-intensive, combining embedding, retrieval, reranking, and large language model (LLM) generation.

By Zhiyuan Cheng, Longying Lai
llmsragefficiencybenchmarkssafety
More like this →
arXiv Machine Learning
Jun 9

AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis

arXiv:2606. 09682v1 Announce Type: new Abstract: AutoMegaKernel (AMK) compiles a HuggingFace Llama-family model into a single persistent cooperative CUDA kernel that runs the whole forward pass in one launch, with no per-model hand-written CUDA.

By Jaber Jaber, Osama Jaber
llmsagentssafety
More like this →
Simon Willison
Aug 12

DeepSeek V4 Pro 0813 (on OpenRouter)

DeepSeek V4 Pro 0813 (on OpenRouter) The latest DeepSeek Pro model is now available, via API only. I had to link to OpenRouter because DeepSeek don't have any obvious announcement page for their new model.

llmsbenchmarks
More like this →
arXiv AI
Jul 28

CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion

arXiv:2607. 23159v1 Announce Type: new Abstract: Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more.

By Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik
diffusionefficiency
More like this →
arXiv Machine Learning
Jun 24

RoPE-Aware Bit Allocation for KV-Cache Quantization

arXiv:2606. 24033v1 Announce Type: new Abstract: Existing low-bit KV-cache quantizers often treat each cached key as a flat vector.

By Fengfeng Liang, Yuechen Zhang, Jiaya Jia
llmsefficiency
More like this →
arXiv Machine Learning
Jun 30

How Far Do On-Prem Open LLMs Get on Text-to-SQL? A Cross-Family Size x Technique Frontier on BIRD

arXiv:2606. 29733v1 Announce Type: cross Abstract: Organizations that cannot send data to a cloud API increasingly ask: how good is Text-to-SQL if the model must run on-premises on open weights, and which popular accuracy "recipes" are worth their compute?

By Vladimir Beskorovainyi
llmsragbenchmarksdata-engineering
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 · bb4ee0e