A paper published on Apple Machine Learning Research and on arXiv compares the best open-source harnesses for autonomous ML engineering with a single coding agent that has only read, write and bash. Under an equal time budget and the same model, the scaffolding gives no advantage. The authors conclude that hand-crafted harnesses around strong models yield poor returns.
- The authors are Kirill Brilliantov, Alejandro Hernández-Cano and Emmanuel Abbé of EPFL; Hernández-Cano did the work while at Apple.
- The conclusion is about current ML engineering benchmarks: orchestrators and retrieval subagents become redundant in the coding agent setting.
- The same week, the industry runs the other way: ever more orchestration around the model.
Read, write, bash. Three primitives and one session. That is the whole "agent" the authors field against the most elaborate open-source systems for autonomous ML engineering, with orchestrators and subagents. And it does not lose.
Hold on before you delete your orchestrator. The conclusion is fenced twice: "current MLE benchmarks" and "equal time budget". This is autonomous ML engineering, a task with a clear result and clear measurement. Your agent that writes proposals and talks to clients is not in this sample, and the paper claims nothing about it.
But the direction hurts. A whole industry is building scaffolding right now: subagents, memory, routers, reviewers, and this week brought more of them. And here comes a paper, via Apple, saying: with a strong model this is weight, not wings. The layers were invented to protect against a weak model. The model is no longer weak, and the layers stayed.
What we do with this: every layer in our agents has to prove it delivers on the same time budget, against the bare session with the same model. If it cannot, it goes. The paper gives the method, and the numbers for your task you pull yourself, because nobody else will pull them for you.