InfoQ Homepage Machine Learning Content on InfoQ
-
From S3 to GPU in One Copy: Rethinking Data Loading for ML Training
Onur Satici discusses Vortex, an open-source columnar file format designed to bypass CPU bottlenecks, enabling ultra-fast S3-to-GPU data streaming and dynamic query pruning at up to 60 Gbps.
-
Running AI at the Edge: Running Real Workloads Directly in the Browser
James Hall explains why engineering teams should shift AI workloads on-device. He shares local inference strategies, WebGPU optimization techniques, and architecture choices for zero-trust privacy.
-
SafeChat: Building AI-Powered Safety Systems at Scale in a Real-Time Marketplace
Bruna Pereira shares how DoorDash built a scalable AI moderation platform. Learn how combining cheap classifiers with LLM scoring reduced incidents and cut latency in real-time chat.
-
From Fab To Token - The State Of The Market
Jordan Nanos explains how hardware constraints, data center scale, and chip co-design shape modern AI performance and tokenomics from silicon to inference.
-
From Models to Agents: Building Context-Aware Consumer AI at Scale at DoorDash
Sudeep Das explains how DoorDash transforms e-commerce recommendations using LLM-driven semantic consumer memory, hierarchical semantic IDs, grounded agentic search, and multi-tiered LLM ranking.
-
The Infrastructure Challenge behind Production AI
The panelists share how core data systems and cloud platforms handle machine-driven workloads, exploring infrastructure design patterns and failure points as AI becomes a permanent fixture.
-
Reimagining Platform Engagement with Graph Neural Networks
Mariia Bulycheva explains how Zalando utilized Graph Neural Networks to move beyond tabular data, modeling higher-order user-item interactions to optimize long-term engagement and click probability.
-
Building Embedding Models for Large-Scale Real-World Applications
Sahil Dua explains the architecture and training of embedding models. He shares practical tips for distilling large models and scaling RAG applications for real-time production environments.
-
How to Unlock Insights and Enable Discovery within Petabytes of Autonomous Driving Data
Kyra Mozley explains Perception 2.0, shifting from rigid CV pipelines to semantic embeddings. She shares how Wayve uses foundation models & vector search to solve the edge case "needle in a haystack."
-
Humans in the Loop: Engineering Leadership in a Chaotic Industry
Michelle Brush discusses engineering leadership in the age of AI/ML and automation.
-
Growing and Cultivating Strong Machine Learning Engineers
Vivek Gupta explains how to nourish and cultivate Machine Learning engineers, detailing the unique production-ML skills required for scaling, governance, and LLMOps.
-
Achieving Precision in AI: Retrieving the Right Data Using AI Agents
Adi Polak discusses achieving precision in GenAI by moving beyond RAG to Agentic RAG. She details agent patterns, feedback loops, and using data streaming architectures to scale real-time AI.