Skip to content

// team structure

In-house ML team vs ML agency

The framing that causes the most damage here is treating this as permanent and exclusive. Most companies that get it right do both, in sequence: contract the first build, hire the owner once there is something worth owning.

The real variables are how many models you will end up running, whether the capability is core to your product, and how much risk you can carry while a hire ramps up.

Side by side

Nine dimensions that decide it.

DimensionIn-house ML teamML agency
Time to first model in productionThree to nine months including search, notice period, and ramp-up.Weeks. The team already has the pipeline patterns.
Annual costLoaded salary plus recruiting, tooling, and management overhead - for the full year regardless of workload.Project cost for the build, optional retainer after. Scales with actual work.
Bus factorA single hire is a bus factor of one until you hire a second.Team coverage during the engagement, then documentation after handover.
Domain knowledge of your businessDeep, and compounds every month they stay.Learned during the engagement. Real but shallower than a year of being inside the company.
Breadth of pattern exposureWhatever that individual has seen before.Patterns from many builds across different data shapes.
Long-term ownershipNatural. Someone is accountable by default.Has to be deliberately arranged at handover, or it degrades unowned.
Retention riskML engineers move often. Losing the only one is severe.Contractual. The engagement ends on a known date with a known deliverable.
Fit for a one-off modelPoor. A permanent hire for a bounded project is expensive.Good. This is the shape the model fits best.
Fit for a model portfolioGood. Fixed cost amortises across many models.Gets expensive past a certain number of concurrent systems.
How to decide

Pick by situation, not by preference.

Hire in-house when

  • ML is core to your product rather than an internal efficiency tool.
  • You expect a portfolio of models, not one or two.
  • You already have a senior technical leader who can hire and evaluate ML people credibly.
  • The domain is unusual enough that context takes months to acquire and keeps paying back.
  • You can survive the three-to-nine month gap before the first model ships.

Use an agency when

  • You need the first model in production this quarter, not next year.
  • The scope is one or two well-defined systems.
  • You are not yet sure ML will pay off here and want the answer before committing headcount.
  • You have no one internally who can evaluate an ML hire, and would be interviewing blind.
  • You want the option to bring it in-house later against a working system rather than a blank page.
Our bias, stated

The hybrid beats both in most cases we see. Contract the first build, get it into production, and use the working system as the thing your first hire takes ownership of. Interviewing for a role where the deliverable already exists is a completely different conversation from interviewing for a role where nobody in the building can assess the answers. It is also why we treat the retraining runbook and handover documentation as part of the deliverable rather than an upsell - a client who takes the system in-house cleanly is a reference, and a client stuck depending on us is a liability for both sides.

Questions

Follow-up questions.

Is a fractional or part-time ML hire a third option?
It can work for maintaining an existing system - watching drift, running scheduled retraining, handling small changes. It works poorly for the initial build, where the work is concentrated and needs continuity. Fractional after a build, not instead of one.
What should we insist on in an agency contract?
Full IP transfer on final payment, the feature pipeline and training code as deliverables rather than the model artefact alone, a written retraining runbook, and no runtime dependency on the agency's infrastructure. If any of those is missing, you are renting a capability, not buying one.
How do we evaluate an agency with no ML person on staff?
Ask how they would define the label for your specific problem and what would make them stop the project. A team that answers with a real definitional discussion and a genuine off-ramp is doing engineering. A team that answers with a tool list is selling integration work.

Still not sure which side you are on?

Thirty minutes, no pitch. We will tell you which of the two your situation actually points to - including when the answer is the one we do not get paid for.

Book a 30-min call
All comparisons