Optimize and deploy with Optimum-Intel and OpenVINO GenAI
Related stories
Accelerating engineering cycles 20% with OpenAI
Accelerating engineering cycles 20% with OpenAI.
Accelerating PyTorch distributed fine-tuning with Intel technologies
OpenAI and Broadcom announce strategic collaboration to deploy 10 gigawatts of OpenAI-designed AI accelerators
OpenAI and Broadcom announce a multi-year partnership to deploy 10 gigawatts of OpenAI-designed AI accelerators, co-developing next-generation systems and Ethernet solutions to power scalable, energy-efficient AI infrastructure by 2029.
Introducing OpenAI o3 and o4-mini
Our smartest and most capable models to date with full tool access
Accelerating Qwen3-8B Agent on Intel® Core™ Ultra with Depth-Pruned Draft Models
Optimizing your LLM in production
Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT
arXiv:2601. 20408v2 Announce Type: replace-cross Abstract: Enterprise LLM deployment faces a critical scalability challenge: organizations must optimize models systematically to scale AI initiatives within constrained compute budgets, yet the specialized expertise required for manual optimization remains a niche and scarce skillset.
GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads
arXiv:2607. 02518v1 Announce Type: cross Abstract: OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the context window.
OpenAI o3-mini
InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
arXiv:2607. 20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or narrow action spaces.
Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
arXiv:2608. 15693v1 Announce Type: new Abstract: Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation.