InfoQ Homepage Performance Content on InfoQ
-
Keeping ChatGPT Fast as AI Development Accelerates
Martin Spier shares how agentic coding accelerates release velocity at OpenAI and explains how always-on AI agents redefine performance engineering to keep ChatGPT fast at massive scale.
-
From ms to µs: OSS Valkey Architecture Patterns for Modern AI
Dumanshu Goyal explains how moving from proxied architectures to direct access in Redis/Valkey slashes tail latency to microseconds, lowers costs, and eliminates single points of failure.
-
Automatically Retrofitting JIT Compilers
Laurence Tratt explains yk, a meta-tracing JIT compiler technology that automatically speeds up C-based interpreters like Lua and Python with minimal, non-invasive code changes.
-
The Rust High Performance Talk You Did Not Expect
Ruth Linehan discusses Momento’s 3-year migration from Kotlin to Rust, explaining how the borrow checker speeds up delivery velocity and why Rust lowers overall engineering costs in production.
-
The Infrastructure Challenge behind Production AI
The panelists share how core data systems and cloud platforms handle machine-driven workloads, exploring infrastructure design patterns and failure points as AI becomes a permanent fixture.
-
Realtime and Batch Processing of GPU Workloads
Joseph Stein explains how to build a highly available private AI cloud. He shares blueprints on scaling vLLM on enterprise GPUs, implementing gateway guardrails, and optimizing batch workloads.
-
The AI Gateway: Scaling Centralized Inference across Decentralized Teams
Meryem Arik explains how AI model gateways resolve the chaos of decentralized engineering teams by centralizing inference. Learn to optimize costs, enforce governance, and maximize GPU utilization.
-
Using AI as a Thinking Partner for Large-Scale Engineering Systems
Google senior staff engineer Julie Qiu shares how she uses AI as a thinking partner to navigate large-scale systems, moving beyond code generation to architecting complex, multi-language ecosystems.
-
The Human Scalability Problem: Why Your Teams Don’t Scale Like Your Code
Charlotte de Jong Schouwenburg explains why scaling engineering teams often slows down delivery. She shares how to solve human latency by building trust and psychological safety across silos.
-
Speed at Scale: Optimizing the Largest CX Platform out There
Matheus Albuquerque explains how to modernize legacy React codebases using jscodeshift, code splitting, and Preact to achieve a 37% bundle reduction while maintaining support for older browsers.
-
Building Embedding Models for Large-Scale Real-World Applications
Sahil Dua explains the architecture and training of embedding models. He shares practical tips for distilling large models and scaling RAG applications for real-time production environments.
-
Scaling to 100+ as a Director: Lessons from Growing Engineering Organizations
Thiago Ghisi discusses his journey scaling engineering orgs to 100+ people. He shares a 3-level impact framework and explains how to evolve from a "solver" to a "driver" of organizational growth.