Theory Ventures is backing Ollama, the most popular way for developers to access open models, which boasts a large developer network. The role involves building a scalable inference platform for developers to run large open models, focusing on high-throughput and low-latency distributed systems.
Responsibilities:
- Build and scale the inference platform that serves every request from ollama.com
- Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability
- Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering
- Build the reliability, observability, and cost controls for our team and customers