Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware—a position that took years to build.

Member of Technical Staff, Inference

Member of Technical Staff, Cloud Orchestration (Remote)

Member of Technical Staff, Cluster Administration

Member of Technical Staff, Site Reliability Engineer

Founding Product Designer

Head of Engineering

Product Marketing Manager

Member of Technical Staff, AMD GPU Performance Engineering

Member of Technical Staff, TPU Performance Engineering

Member of Technical Staff, AMD GPU Performance Engineering

Member of Technical Staff, Performance and Scale















