Together AI — Inference Engineer
Post-sales engineering for dedicated inference endpoints.
Together AI · dedicated endpoints & CX engineering · B200 / GB300 / H200 / H100 fleets
My customers run open-weight frontier models — Kimi K2.6 / K3, MiniMax M2.7 / M3, and other
large MoE, long-context, and multimodal models — as single-tenant deployments where an outage
is their product going down. I own that surface end to end, working hands-on across
TensorRT-LLM, vLLM, and SGLang, down to engine source when needed.
Ownership
The revenue-critical edge
Primary engineering owner for 3 of the platform's top-10 revenue accounts ($8M+ quarterly GMV): a frontier AI lab, one of the largest AI-coding products, and enterprise conversational-AI customers.
Depth
Root causes, not ticket ping-pong
When a flagship deployment kept failing in ways restarts couldn't fix, I read the serving-engine source, isolated the failure, found the upstream open-source fix, and authored the PR.
Outcomes
Ship the fix and the story
Caught an API parameter that was accepted but silently ignored across service layers; wrote the engineering handoff that unblocked a customer's production launch.
Serving stacks
Disaggregated serving, at full tilt
Support NVIDIA Dynamo deployments in production for frontier MoEs (Kimi K2.6, K3) — prefill/decode disaggregation and KV-aware routing that pull full capacity out of B200-class fleets.
Automation
Days to minutes, literally
Built AI-powered triage skills — a diagnostic toolkit that walks the full request path from edge to engine — cutting self-serve customer resolution time from days to minutes.
Observability
See the problem before the page
Ship Grafana dashboards and alerts as deployments need them, plus a generic alerting framework — and new diagnostic methods that pinpoint what's wrong with a model or deployment in minutes, feeding findings back into product fixes.