About this session
How do you serve 168 models through one integration and stay reliable at 3M+ requests a day? Adapta, the team behind Brazil's largest generative AI platform, did exactly that on AI Gateway and the Vercel Platform. Hear how they built an app-level router that picks the right model per request in under 200ms, wired up automatic failover and per-user budgets, and kept reliability climbing as traffic spiked. A practical look at the architecture decisions behind running frontier models at scale.
What you’ll take away
- Compare the tradeoffs of application-level model routing.
- Understand how failover and per-user budgets support reliability.
- Examine an architecture serving millions of requests a day.
Speaker

Max Peters
CEO · Adapta
Max Peters is CEO of Adapta. At Ship, he shares how the team routes requests across a large model catalog with AI Gateway, covering automatic failover, per-user budgets, and the architecture choices that keep a high-traffic AI platform reliable.