InfoQ Homepage application performance management Content on InfoQ
-
Multi-Agent Patterns from Spotify’s AI Powered Advertising Platform
Pratik Rasam explains how Spotify Ads Manager builds multi-agent AI systems in production using Google ADK Java, detailing guardrails, domain ownership, and key architectural patterns for scale.
-
Building Reusable Evaluation Frameworks for Agentic AI Products
Susan Chang shares how Elastic built a production-grade AI agent evaluation framework, detailing lessons on tracing, LLM-as-a-judge, programmatic evals to prevent regression across engineering teams.
-
Beyond Observability: Evolving Production Operations in the Age of AI
The panelists share how AI and automation reshape production operations, turn system data into actionable insights, and change architectural practices for modern, complex software delivery.
-
Context is the New Code
Patrick Debois explains how treating context like code transforms the Software Development Life Cycle into a Context Development Life Cycle to scale AI agent workflows reliably across teams.
-
Adaptive Recommenders in the Real World: Inference, Evals, and System Design
Mallika Rao shares why building adaptive recommendation systems requires shifting focus from isolated ML models to real-time feedback loops, retrieval freshness, and production-level constraints.
-
From AI Agent Demo to Production: Automated Testing and Evaluation
Zhou Yu explains why 95% of AI agents fail to reach production and shares how simulation-driven evaluation, synthetic users, and automated CI/CD testing unlock enterprise deployment.
-
Can Claude Fix Itself? Using LLMs for Incident Response
Anthropic's Alex Palcuie shares how LLMs transform incident response, highlighting where Claude excels at log analysis and why automated AI SREs still can't replace human judgment.
-
Enchant Your AI and APIs with eBPF Magic 🪄
Dan Finneran explains how to use eBPF and AI gateways in Kubernetes to transparently observe, modify, and control unowned AI agent API calls without altering source code.
-
Keeping ChatGPT Fast as AI Development Accelerates
Martin Spier shares how agentic coding accelerates release velocity at OpenAI and explains how always-on AI agents redefine performance engineering to keep ChatGPT fast at massive scale.
-
From Confusion to Clarity: Advanced Observability Strategies for Media Workflows at Netflix
Naveen Mareddy and Sujana Sooreddy explain how Netflix monitors massive media encoding workflows. They discuss scaling to 1M+ trace spans and using Flink to unlock real-time business insights.
-
From Dashboard Soup to Observability Lasagna: Building Better Layers
Martha Lambert shares the "Observability Lasagna" framework, explaining how her team replaced "dashboard soup" with a user-centric stack to achieve high system reliability for their on-call product.
-
Scaling API Independence: Mocking, Contract Testing & Observability in Large Microservices Environments
Tom Akehurst discusses how API mocking and simulation can solve microservice decoupling and productivity problems at scale, using observability, contract testing, and GenAI as guardrails.