Talk it through with Aurelius
Library›Aurelius›The problem
Aurelius · Work & Leadership
Knowledge + Guidance

We can train the model. Why does the system fall apart?

You built something that works on your machine. Then it met the world — other people's code, other people's timelines, a pipeline that has to run at three in the morning without you watching it — and it did not survive contact. You call this a failure of skill. It is not. No one told you that ML engineering and ML systems engineering are separate crafts. One is about the model. The other is about everything the model depends on to stay useful after you stop touching it. Your pod has people who can do the first. Almost no one has clearly claimed the second. That is the actual hole. You cannot fix a job no one has named. So name it this week, in front of your pod, out loud. Then choose who owns it. Not the most talented person — the one willing to hold it.

◆ How this problem reads on the two dials
GuidanceKnowledge
Coaching
More to learn
1:1 with AureliusWith others (a Pod)
Some one-to-one
Practise with peers
The pod needs to first understand that model-building and systems-building are distinct disciplines with different skills — a teaching gap — before coaching on division of labor will hold.
How the two dials adapt to you →
What’s really going on

Because training a model and building a system are two different jobs, not one. The first is solitary craft. The second is coordination — pipelines, monitoring, handoffs, people. Your pod is failing at a job no one named, not at the job you already know. Name it. Assign it. Then work.

🔒 What you’ll build togetherUnlock by starting
A moveList every task the model touches after training. Split the list into two columns: 'model' and 'system.' Do this together, out loud, this week.
A moveAssign one person to own the system column. Not the most talented — the one willing to be accountable for it.
A moveStop calling deployment failures 'bugs in the model.' Name them precisely: data drift, handoff failure, monitoring gap, no owner.
A movePick one system task on your list. Finish it before anyone starts a new modeling experiment.
A moveEach week, ask the pod one question: what broke, and was it the model or the scaffolding around it? Write down the answer.
PractiseMap the Seams · a Pod of 4 · 30 min

What changes unlock by starting

  • Your pod names ML engineering and ML systems as two separate jobs, and stops pretending one person or one skill covers both.
  • Someone owns system health specifically, so it is not everyone's job and therefore no one's.
  • You stop diagnosing systems failures as modeling failures, and waste less time retraining what was never broken.
  • Fewer surprises reach production, because someone is watching the seams instead of only the model.
One object, two jobs: a public answer to a real problem, and — the moment you start the chat — Aurelius’s live plan for your version of it.