Sam Austin on October 6, 2026

Building an ML Platform Team: Roles, Tools, and Org Structure

Building an ML Platform Team: Roles, Tools, and Org Structure
Contents

Data and ML platform team collaborating on production infrastructure

Figure 1: The platform team is an org design problem before it is a technical one — the right structure lets model builders ship without stepping on each other

Your data scientists are brilliant and your models work great in notebooks. Then someone asks "can we deploy this by Friday" and the whole thing falls apart, because nobody owns the path from trained model to production endpoint. That gap is exactly what an ML platform team exists to close, and building one well is as much an org design problem as a technical one — the technical side is covered in the MLOps beginners guide; this is the people side.

Let's work through the roles that actually matter, the structures that scale, and the tooling landscape as it stands today.

Why This Is an Org Design Problem First

The instinct when ML deployments feel chaotic is to hire more engineers. The better first move is figuring out what's actually broken: is it that nobody owns deployment infrastructure, or that everyone sort of owns it and therefore nobody does? MLOps team structure differs fundamentally from regular software engineering because it genuinely involves three distinct skill sets that rarely live in one person: data engineering (pipelines and feature stores), ML engineering (model training and experimentation), and platform engineering (serving, infrastructure, and monitoring). Most dysfunction traces back to one of these three being unowned, not to a lack of raw headcount.

The Core Roles, and What Each One Actually Owns

Role titles in this space are genuinely not standardized — the same responsibilities show up as "AI Platform Engineer," "ML Infra Engineer," "AI Reliability Engineer," and half a dozen other labels depending on the company you're looking at. That said, a working platform team tends to need some combination of these functions, regardless of what the org chart calls them.

The ML Engineer turns research into shippable, tested, maintainable model code — the bridge between a data scientist's notebook and something that can actually run reliably in production.

The MLOps Engineer industrializes the models once they exist: building training pipelines, automated deployment systems, production monitoring, and model governance. A genuinely useful way to think about the distinction that trips up most hiring managers: keeping models reliable on their way to and through production is MLOps, building an internal platform other teams ship on is ML platform engineering, and building the models themselves is ML engineering. Those are three different jobs that frequently get bundled into one job posting, which is exactly why hiring for this space is so confusing from the outside.

The ML or Platform Architect sets the technical direction, deciding build-versus-buy on major platform components and making sure the pieces different teams are building actually fit together.

The Data Engineer owns the pipelines and feature stores that feed everything downstream — arguably the least glamorous role and also the one whose failures are most visible everywhere else.

The Platform or Infrastructure Engineer owns compute, GPUs, cost, and the security of the whole stack — the layer beneath MLOps that most data scientists never think about until it breaks.

A Data Scientist rounds out the group as the researcher, the person actually building and evaluating models, whose work the rest of the team exists to operationalize.

A Genuinely Useful Hiring Heuristic

Teams light on this problem default to asking "when do we need an MLOps engineer?" The more useful framing from recent hiring guidance: hire an MLOps Engineer when the organization is ready to move models into production at scale, and hire an ML Engineer when the organization is in active model development. If you're still iterating heavily on model architecture and haven't shipped much, more MLOps headcount won't fix your actual bottleneck. If you've got working models stuck in notebooks because nobody can reliably deploy them, that's exactly the MLOps gap.

Three Org Structures, and When Each Makes Sense

Fully embedded, where ML engineers sit directly inside product teams with no central platform function, works fine at early stage — roughly one to three generalist ML or data science hires with part-time data engineering support. Past that size, this model starts producing exactly the chaos that prompts most teams to build a platform function in the first place: every product team reinvents deployment tooling, nobody standardizes monitoring, and model governance becomes inconsistent across the organization.

Fully centralized, a single MLOps team serving every product team's deployment needs, solves the duplication problem but introduces a different failure mode. A centralized team focused on stability and standardization often doesn't have the latest tooling a specific product team needs immediately available, or worse, treats a product team's preferred new tool as "not production-ready" and blocks it. This is where centralized MLOps starts feeling like a toll booth instead of a service — every deployment request queued behind a team that doesn't share the requesting team's urgency or context.

Hub-and-spoke is the model that tends to work best as companies grow past the early stage, and it's worth understanding why. The hub is a central MLOps or platform team owning tooling, deployment infrastructure, monitoring, and guardrails; the spokes are ML engineers and data scientists embedded directly in product teams, using the hub's tooling rather than each reinventing it. This gets you standardization without the toll-booth problem, since spoke teams can move at product speed while still building on shared, maintained infrastructure rather than bespoke one-off pipelines.

A reasonable staged progression: early stage runs one generalist ML engineer, one data engineer, one data scientist, and part-time MLOps coverage. Growth stage brings five to fifteen ML or data science hires alongside two to five data and platform engineers, with a core ML platform team starting to emerge as its own distinct function. Late stage sees specialized ML teams organized by product line, a formal ML platform org, and dedicated governance and risk management roles — the kind of structure regulatory requirements increasingly demand at scale.

How Many Platform Engineers Do You Actually Need

This is the question every leader asks once the team exists, and there's a useful rough benchmark here: plan on roughly one MLOps or platform engineer for every four to six people actually building models. Push that ratio thinner and deployment becomes the bottleneck, with expensive, specialized modelers sitting idle while a single overloaded platform engineer drowns in deployment requests from every direction.

Zooming out further, across a benchmark of hundreds of data teams, the median data team runs at roughly 13% the size of the broader engineering org, with fintech running heavier at around 3.5% of total headcount and B2B running lighter at around 2.4%. These numbers are directional, not prescriptive — your actual right-sizing depends heavily on how many models you're running in production and how critical they are, but they're a reasonable sanity check against wildly under- or over-staffing relative to peers.

Reporting Lines

An AI or ML platform team typically groups product, research, ML engineering, MLOps, data engineering, and infrastructure functions under a single owner, usually a Head of AI or a VP reporting to the CTO. At the individual level, an MLOps Engineer commonly reports to a Lead ML, a Head of Data, or a Head of Platform, and works across three distinct partnerships: with data scientists and ML engineers to productionize their models, with data engineers on upstream pipelines, and with platform engineers on the underlying infrastructure.

Career progression from an MLOps Engineer role typically runs toward Lead MLOps (acting as technical reference for the broader ML platform team), Head of ML Platform or Head of AI Platform (setting strategy and roadmap at scale), or Staff Engineer ML Platform (providing cross-functional technical depth in larger organizations). Some engineers pivot toward a more generalist Platform Engineer or SRE track instead, and some move into ML Solutions Architect roles at major cloud or AI vendors.

Enforcing Structure Through Your Repository, Not Just Your Org Chart

One pattern worth adopting regardless of team size: let your repository structure mirror team ownership directly, rather than relying purely on documentation or tribal knowledge about who owns what. A directory layout where data engineers own /data, ML engineers own /models, and platform engineers own /serving, enforced through CODEOWNERS rules requiring the right team's review on pull requests touching their directory, turns org structure into something Git actually enforces rather than something that exists only in a wiki page nobody reads.

Pair this with a genuine handoff drill early on — a data engineer pushes dummy data, an ML engineer trains on it, a platform engineer deploys the result — to confirm the handoffs between roles actually work before you're depending on them during a real incident. The deployment half of that drill is exactly what the GitOps for ML setup automates once the team exists.

The Tooling Landscape Mapped to Roles

Rather than listing every MLOps tool that exists, it's more useful to map tooling to the specific role responsible for it, since tool sprawl usually happens when nobody's clearly accountable for a given layer.

Data engineers typically own the feature store and data pipeline tooling — Feast or Tecton for feature serving, Airflow or Dagster for orchestration, plus whatever data versioning tool (DVC is common) ties data snapshots to specific model training runs.

ML engineers live primarily in experiment tracking and model development tooling — MLflow or Weights & Biases for tracking runs and comparing results, plus whatever training framework and distributed training infrastructure the team has standardized on.

MLOps and platform engineers own the deployment and serving layer — Kubernetes for orchestration, a model registry (often MLflow's registry component or a cloud-native equivalent) connecting trained artifacts to deployment status, CI/CD tooling adapted for ML-specific needs like model validation gates, and monitoring stacks (Prometheus, Grafana, plus ML-specific drift and performance monitoring) watching models in production. The CI/CD for ML pipeline is this layer end-to-end, and the MLOps platform comparisons show how vendors bundle these pieces.

Security and compliance, increasingly its own dedicated role as organizations hit regulatory requirements like the EU AI Act, typically owns policy-as-code tooling, audit trail automation, and the governance layer tracking which models are approved for which use cases.

Practical Signals You Need a Dedicated Platform Function

A few concrete signs it's time to formalize an ML platform team rather than continuing with embedded, ad hoc deployment:

  • Multiple product teams are independently solving the same deployment problem with different tools.
  • Nobody can answer "what's currently running in production and who owns it" without a multi-person Slack thread.
  • Model governance and compliance requirements are starting to apply across the organization rather than team by team.
  • Your best ML engineers are spending more time on deployment infrastructure than on actual modeling work.

Any one of these is a reasonable trigger; two or more together is a clear signal the ad hoc approach has run its course.

Common Mistakes to Avoid

  • Hiring an MLOps engineer before you have models worth productionizing wastes a specialized, expensive hire on a problem you don't have yet — hire ML engineers first if you're still in active model development.
  • Fully centralizing deployment without a spoke model creates the toll-booth problem, where product teams wait on a platform team with different priorities and less urgency around their specific launch.
  • Letting role titles substitute for actual job clarity — bundling model development, deployment, and infrastructure under one vague title leads to hiring mismatches and unclear accountability once something breaks in production.
  • Skipping the repository-level ownership structure and relying purely on documentation means org structure exists on paper but doesn't actually govern who reviews what, which quietly erodes over time as the team grows.
CoverBookDescriptionGet it
Cover of “Designing Machine Learning Systems” Designing Machine Learning Systemsby Chip Huyen covers the platform thinking behind this article: ownership of data pipelines, deployment, and monitoring as one system rather than three disconnected jobs. View on Amazon
Cover of “Building Data Science Teams” Building Data Science Teamsby D.J. Patil the classic on exactly this question: how to staff, structure, and grow the people side of a data and ML organization. View on Amazon
Cover of “Machine Learning Engineering” Machine Learning Engineeringby Andriy Burkov a practitioner's reference for the MLOps lifecycle that clarifies where each role's responsibility begins and ends. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is an ML platform team?

An ML platform team is the function that owns the path from trained model to production endpoint: training pipelines, deployment infrastructure, model registries, monitoring, and guardrails. It closes the gap between models that work in notebooks and models that run reliably in production, so data scientists and ML engineers can ship without each product team reinventing deployment tooling.

What is the difference between MLOps, ML platform engineering, and ML engineering?

Keeping models reliable on their way to and through production is MLOps; building an internal platform other teams ship on is ML platform engineering; building the models themselves is ML engineering. Three different jobs that frequently get bundled into one job posting, which is exactly why hiring for this space is so confusing from the outside.

When should you hire MLOps engineers versus ML engineers?

Hire an MLOps engineer when the organization is ready to move models into production at scale, and hire an ML engineer when the organization is in active model development. If you are still iterating on model architecture, more MLOps headcount will not fix your bottleneck; if working models are stuck in notebooks because nobody can reliably deploy them, that is the MLOps gap.

How many platform engineers do you need per model builders?

Plan on roughly one MLOps or platform engineer for every four to six people actually building models. Push that ratio thinner and deployment becomes the bottleneck, with expensive specialized modelers sitting idle while a single overloaded platform engineer drowns in deployment requests from every direction.

Wrapping Up

A working ML platform team isn't really about hiring more engineers — it's about clearly assigning the three distinct skill sets (data, ML, and platform engineering) that production ML genuinely requires, and choosing an org structure — embedded, centralized, or hub-and-spoke — that matches your actual scale rather than copying whatever structure a much bigger company uses.

Will you need every role covered here on day one? No — start with ML engineering if you're still in active model development, and add dedicated MLOps capacity once you've got models genuinely worth productionizing at scale. But get the repository ownership and reporting lines explicit early, since retrofitting clear accountability onto a team that's already grown chaotic is a much harder problem than building it in from the start.

What are You Looking For?

esc