Technical skills
Methods
Role signals
About StackYak
StackYak is building the infrastructure layer for AI.
We are an early-stage, funded company building software that brings compute, GPU infrastructure, networking, and inference together into one product. The opportunity is large, the market is moving quickly, and we are building for real production workloads from the start.
This is not internal IT. This is not a slow-moving infrastructure team maintaining someone else's platform. The infrastructure is the product .
We are a small, senior team with very little bureaucracy. This is a founding role in its discipline. You will be the first person here whose primary responsibility is the systems layer beneath the product, and the shape it takes will largely be yours to decide.
Treat this document as a starting point rather than a boundary. The people who do well here take ground early and are not asked to give it back.
We move quickly. We do not have months for someone to learn the fundamentals of their discipline. You should already be very good at what you do, be able to ramp into adjacent areas quickly, and be comfortable operating without perfect requirements or neatly defined boundaries.
The Role
We need someone who can take a model, a set of GPUs, and a production requirement and determine how that model should actually run .
Hand you a frontier open-weight model — dense or mixture-of-experts, tens to hundreds of billions of parameters — along with the hardware actually available and the latency a customer actually needs, and you should be able to make the calls that follow: precision, quantization, tensor and pipeline and expert parallelism, memory strategy, batching, concurrency, topology, and serving runtime. Then defend them, measure them, and operate them.
Doing that well is the baseline. What makes the role interesting is the question sitting underneath it.
Serving expertise is scarce, slow to acquire, and currently lives in the heads of a small number of people who have done it enough times to have the instinct. We do not accept that it has to stay that way. A great deal of what those people know is reasoning, and reasoning can be written down, tested, and eventually carried out by something other than a person at two in the morning. How far that can be pushed is genuinely open, and you would be one of the people finding out.
This is not a research position and it is not an architecture-only role. You will benchmark, deploy, debug, tune, automate, and operate real inference systems — and then carry them in production.
What You Will Own
Not tasks. Outcomes, and the authority that comes with them.
- How models run. Placement, parallelism, precision, and memory strategy across single-GPU, multi-GPU, and multi-node — and the reasoning behind each choice, in a form someone other than you can audit.
- The serving runtimes. vLLM, SGLang, TensorRT-LLM, and whatever displaces them: which one we use where, and what we do when the one we picked turns out to be wrong.
- What we are willing to promise. Concurrency limits, latency and TTFT targets, throughput under load. You own the evidence behind a number before it becomes a commitment to a customer.
- Benchmarking as an institution. Methodology, harness, reproducibility — and the standing to say that a result does not mean what someone wants it to mean.
- Model lifecycle in production. Loading, startup, health, upgrades, capacity, and failure handling, all of it happening while customers are connected.
- Multi-tenant versus dedicated serving , and where the line between them actually falls.
- Inference incidents , including the ones where the serving runtime turns out to be innocent and the cause is GPU, driver, host, or network.
- Turning all of the above into software , so that it stops being institutional memory with a single point of failure.
What Success Looks Like
You can be given a model, a hardware target, and a workload profile and quickly produce a deployment that is sensible, measurable, reproducible, and ready to operate.
Then you help us make that expertise programmatic.
You will help us answer questions such as:
- What hardware can this model run on?
- What is the right precision or quantization strategy?
- How should we split it across GPUs or nodes?
- What serving runtime should we use?
- How much concurrency can we safely offer?
- How should we trade latency against throughput for this particular workload?
- How does StackYak make those decisions automatically instead of needing an expert every time?
Requirements
- What We Need
- The bar is what you have already done, not what you could learn. You should have done most of this:
- Run large language models in production, under real load, with someone depending on them.
- Served a model across multiple GPUs, and across multiple nodes.
- Sized a model against hardware and been right about whether it would fit.
- Chosen a quantization and precision strategy with a real quality and cost consequence attached to the choice.
- Tuned batching and concurrency past the point where the easy wins ran out.
- Measured latency, TTFT, throughput, utilization, and cost — and defended the numbers to someone motivated to disbelieve them.
- Worked in NVIDIA and/or AMD inference environments.
- Written Python that other people run in production: tooling and product, not scripts.
- Debugged Linux and systems problems below the container boundary.
- Found the root cause of an inference failure that was not in the serving runtime.
You Will Be Especially Strong If
- You have run inference infrastructure at an AI company, inference provider, GPU cloud, hyperscaler, or serious internal AI platform.
- You have worked across multiple GPU generations and vendors, and can say concretely how that changed your decisions.
- You understand distributed inference and the networking implications of multi-node serving.
- You have benchmarked and compared serving runtimes rather than treating one framework as the answer to every problem.
- You can explain why a deployment is configured the way it is instead of repeating a vendor recipe.
- You have automated deployment decisions, or built schedulers, placement systems, capacity planners, or similar infrastructure.
- You build and test things outside your assigned roadmap because you want to know how they actually work.
This Is Probably Not For You If
- Your background is primarily model training or ML research, with little production serving behind it.
- You have put a model behind an API but have never had to reason about GPU memory, parallelism, topology, and concurrency.
- You treat vLLM defaults as an architecture.
- You are interested in benchmarks but not in operating the systems they describe.
- You want to optimize kernels all day and have little interest in reliability or productization.
- You would rather stay the person who knows how to configure it than encode that knowledge into software that makes you unnecessary.
- You prefer writing recommendations to implementing them.
- You need narrowly defined ownership, or a long runway before taking responsibility.
Benefits
How We Work
- Small, senior team with direct access to the founders.
- Strong opinions, loosely held.
- Everyone is expected to participate in technical decisions.
- Everyone shares responsibility for production and on-call.
- We value people who can move between design, implementation, debugging, and operations.
- We care much more about what you have built and operated than degrees, certifications, or academic credentials.
- We expect people to leave ego at the door, argue the technical case, make a decision, and then execute.
- We are remote and distributed across time zones. Whether a role is an employment or a contract engagement depends on where you are, and we work that out at offer.
- Hiring here is a few real conversations with the people you would actually work with, not a recruiter screen followed by a panel of strangers.
- We are hiring across inference, infrastructure, and networking. The boundaries between the three are blurry on purpose. If you sit between two of them, say so.
- This is an early-stage startup. The pace is high, the problems are hard, and the scope will change as we grow.
Compensation
Competitive compensation plus meaningful equity. Exact structure will depend on location, engagement model, and experience.
A Note For Agencies
We are not using external recruiters or agencies for this role, and we will not be persuaded otherwise by an email. We do not want your spam. We will not read the CVs you send, we will not reply to your follow-up, and no candidate you put in front of us creates a fee obligation of any kind. Do not contact us.