Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models
Related stories
Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B
arXiv:2607. 04244v1 Announce Type: new Abstract: This report describes our approach to the Efficient Qwen Competition, where the goal is to enable low-latency serving of Qwen3.
Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead
arXiv:2608. 13987v1 Announce Type: new Abstract: Nanbeige4.
QAgent: An LLM-based Multi-Agent System for Autonomous OpenQASM programming
arXiv:2508. 20134v2 Announce Type: replace Abstract: Programming quantum circuits at the OpenQASM level is essential for achieving hardware-aware optimization and reliable execution on noisy intermediate-scale quantum (NISQ) devices, yet it remains challenging due to the need for domain-specific planning, iterative code synthesis, and low-level calibration.
Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
Friday's big release was Qwen 3. 8 27B , an Apache 2 licensed 27B parameter vision-capable LLM from Alibaba's Qwen research lab.
MaxKernel: Agentic Kernel Generation for TPUs
MaxKernel is a multi‑agent system designed to generate high‑performance custom kernels for TPUs. It offers three paradigms: a Human‑in‑the‑Loop agent for collaborative design, an Autonomous agent that runs a fully automated optimization loop, and a Graph‑Based Autonomous Search for global exploration. All paradigms share specialized sub‑agents for planning, implementation, debugging, testing, and profiling, and the system consistently matches expert hand‑tuned baselines on the JaxBench suite and real‑world workloads.
UltraQuant: 4-bit KV Caching for Context-Heavy Agents
arXiv:2606. 20474v1 Announce Type: cross Abstract: Context-heavy agents place unusual pressure on the key-value (KV) cache: long prefixes are reused across many short turns, while concurrency determines whether the serving system can keep GPUs utilized.
Can Agents Design Better Chips with a Higher Level Abstraction?
Large Language Model agents are being explored for chip design, but most methods work directly at RTL. This study compares Direct RTL Design, Agent-based HLS Design, Post-Compiler HLS Refinement, and Post-HLS RTL Refinement, and proposes a combined workflow called Agent-based HLS with RTL Refinement (AHRR). Using FPGAs for evaluation, AHRR achieves a 2.6× geometric‑mean speedup over Direct RTL Design across an 11‑task benchmark suite, demonstrating that higher‑level abstractions and subsequent RTL refinement can improve chip design efficiency.
Can I Buy Your KV Cache?
arXiv:2606. 13361v1 Announce Type: new Abstract: Right now, across the world, AI agents are repeating the same absurd act: to read one document, they each recompute it from scratch.
Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
arXiv:2608. 15693v1 Announce Type: new Abstract: Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation.
FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon
arXiv:2607. 22785v1 Announce Type: cross Abstract: Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units.
Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI
arXiv:2608. 26418v1 Announce Type: cross Abstract: Modern AI workloads and the hardware that runs them evolve on different timescales: architectural definition precedes volume silicon by years, while target workloads shift in months.