Overview

TensorX is a sovereign AI infrastructure platform headquartered in Dublin. We run frontier open-weight large language models on our own NVIDIA Blackwell GPUs in European datacentres, under EU jurisdiction. Customers reach them through a drop-in OpenAI-compatible API. Nothing they send is retained after the request completes. We serve regulated enterprises in finance, healthcare and government as well as developers and AI platforms. We help them adopt AI without compromising on data privacy, compliance or performance.


We are looking for a Senior Inference Engineer to join our growing engineering team. Reporting to the CTO, you will own the serving layer between our API gateway and the inference engines across more than one site. This covers admission, routing, the prefill/decode split, KV cache movement across the GPU fabric and autoscaling.


Our customers send very long contexts at high concurrency, often in bursts. The number that decides our margin is not how fast one engine runs. It is how many requests the whole fleet answers inside a latency target when a burst arrives. We run NVIDIA Dynamo in production for disaggregated serving on Kubernetes and are moving the rest of the fleet onto it. We are an AI-native team. Tools such as Claude Code and Codex are part of our daily workflow and materially accelerate how we build and operate systems.


You will work side by side with our GPU Performance Team, who own the inside of the engines: kernels, engine patches and per-model tuning. They make each worker faster. You make the fleet of workers behave. Our platform team builds and maintains the Kubernetes clusters and hosts. Our backend team owns the API gateway. You will work with all three every day.


This is a high-impact senior individual contributor role spanning disaggregated serving, routing, autoscaling and multi-site operations.

Responsibilities

  • Disaggregated serving - Run and extend our NVIDIA Dynamo deployment. Own the split between prefill and decode workers. Move capacity between them as the traffic shape changes.

  • Admission & routing - Own the router and KV-aware routing so requests land where their prefix is already cached. Build admission control so a burst queues instead of crashing a worker. The GPU Performance Team tunes cache behaviour inside the engine and contributes to the router code.

  • Autoscaling & capacity - Scale prefill and decode against latency targets and scale down when traffic drops. Decide how many GPUs each model pool needs and where it runs. Base these decisions on traffic data and the parallelism layouts the GPU Performance Team validates.

  • KV cache transfer - Own cross-node KV cache movement (NIXL, Mooncake or similar). Test it across nodes so the network is part of every measurement. Add custom telemetry where the stock tools do not see.

  • Multi-site serving - Run several sites behind one front door, each with its own local serving plane. Keep a path open to non-NVIDIA pools through llm-d when we need one.

  • Observability - Instrument the serving layer with per-model metrics that mean something: time to first token, inter-token latency, queue depth, cache hit rate and errors. Work with Prometheus, Grafana and Loki alongside GPU telemetry.

  • Rollouts & pre-production - Roll out serving-layer changes with zero downtime: drain, verify outside the pool, return, repeat. Share the pre-production gate with the GPU Performance Team so no change reaches a customer unmeasured.

  • Data retention - Make sure nothing a customer sends is retained after the request completes. Enforce this by design in the serving path, not only by policy.

  • Reliability & incidents - Own the serving-layer runbook. Lead the response when the fleet, rather than one engine, misbehaves. Write up each incident so it does not happen twice.

Skills & Experience

  • 5+ years of professional experience in distributed systems, ML infrastructure or production serving. A meaningful portion of this should be on GPU workloads

  • Hands-on experience running large language models in production with vLLM, SGLang or TensorRT-LLM. You understand how prefill, decode, KV cache and batching behave under load

  • Experience with NVIDIA Dynamo or a comparable disaggregated serving system (e.g. llm-d)

  • Kubernetes in production, including deployments, rollouts, operators or controllers and GPU scheduling

  • Strong distributed systems fundamentals: queueing, backpressure, admission control, routing and retries under burst load

  • Experience instrumenting production systems (e.g. Prometheus, Grafana, Loki) and using that data to guide tuning decisions

  • Familiarity with benchmarking methodology: one variable at a time, a pass or fail bar set before the run and results someone else can reproduce

  • Proficiency in Python and comfort reading systems code in other languages

  • Comfortable using AI-assisted development tools (e.g. Claude Code, Codex) as part of your daily workflow

  • A clear and concise communicator who thrives in ambiguity and can articulate technical decisions to both technical and non-technical audiences

Nice to Have

  • Experience with NIXL, Mooncake, LMCache or other KV cache transfer and offload layers

  • Experience debugging performance across a GPU fabric (RDMA, RoCE or InfiniBand)

  • Gateway API Inference Extension or similar Kubernetes-native inference routing

  • Multi-site or multi-cluster operations

  • Contributions to open-source inference or serving projects

Why This Role

  • Disaggregated serving is live here, not on a roadmap.You will extend a Dynamo deployment that already carries production traffic.

  • The fleet is ours.Our own NVIDIA B300 GPUs in Dublin and Helsinki. You are not renting time on someone else's cluster.

  • The traffic is real.Long-context, bursty, high-concurrency workloads that break assumptions benchmarks never test.

  • The results are real.Our inference stack answers roughly twice as many requests inside the latency target as a standard configuration. We measured this on the same hardware with the same production traffic patterns.

  • You shape the architecture.Dynamo is live, but the multi-site serving layer around it is still being designed. Your decisions become the architecture.

Education & Qualifications

  • BSc/MSc in Computer Science, Software Engineering, Electrical Engineering or a related technical discipline OR equivalent practical experience

Remuneration

  • Highly competitive package, dependent on experience

  • 25 days paid annual leave

  • Hybrid working from our centrally located Dublin office, with remote flexibility

  • Free inference tokens!

******* NO AGENCY ASSISTANCE REQUIRED *******


Apply for position now

Are you currently eligible to work in Ireland or the EU without sponsorship?
Which of these have you used in production? (Select all that apply)
Which best describes your experience with disaggregated serving (separate prefill and decode workers)?
What is your notice period or earliest available start date?