we_are_coded.by CODE · The world, decoded
БГ
Apple

Study at Apple: a strong model with a minimal harness keeps up with elaborate agent scaffolding on an equal time budget

Apple Machine Learning ResearchScience

A paper published on Apple Machine Learning Research and on arXiv compares the best open-source harnesses for autonomous ML engineering with a single coding agent that has only read, write and bash. Under an equal time budget and the same model, the scaffolding gives no advantage. The authors conclude that hand-crafted harnesses around strong models yield poor returns.

In short
  • The authors are Kirill Brilliantov, Alejandro Hernández-Cano and Emmanuel Abbé of EPFL; Hernández-Cano did the work while at Apple.
  • The conclusion is about current ML engineering benchmarks: orchestrators and retrieval subagents become redundant in the coding agent setting.
  • The same week, the industry runs the other way: ever more orchestration around the model.
Checked on2 October 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

Read, write, bash. Three primitives and one session. That is the whole "agent" the authors field against the most elaborate open-source systems for autonomous ML engineering, with orchestrators and subagents. And it does not lose.

The facts: in October 2026 Apple Machine Learning Research published the paper "How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?" by Kirill Brilliantov, Alejandro Hernández-Cano and Emmanuel Abbé (EPFL; Hernández-Cano did the work while at Apple). It was submitted to arXiv on 30 September 2026. The authors compare modern agents for autonomous ML engineering, which sit on increasingly elaborate machinery (multi-agent orchestrators, dedicated retrieval subagents), with a more primitive coding agent in which the model has direct access to the environment through read, write and bash primitives. The finding: under an equal time budget and the same frontier LLM backbone, open-source state-of-the-art harnesses provide no advantages over a single session of a minimal-harness coding agent; the backbone is the primary driver of performance. Through a series of systematic studies that strip the layers away one by one, the authors argue that the machinery layers become redundant in the coding agent setting, and conclude that the effort spent elaborating hand-crafted harnesses around strong models yields poor returns for current MLE benchmarks.

Hold on before you delete your orchestrator. The conclusion is fenced twice: "current MLE benchmarks" and "equal time budget". This is autonomous ML engineering, a task with a clear result and clear measurement. Your agent that writes proposals and talks to clients is not in this sample, and the paper claims nothing about it.

But the direction hurts. A whole industry is building scaffolding right now: subagents, memory, routers, reviewers, and this week brought more of them. And here comes a paper, via Apple, saying: with a strong model this is weight, not wings. The layers were invented to protect against a weak model. The model is no longer weak, and the layers stayed.

The scaffolding was built for a model that no longer exists.

What we do with this: every layer in our agents has to prove it delivers on the same time budget, against the bare session with the same model. If it cannot, it goes. The paper gives the method, and the numbers for your task you pull yourself, because nobody else will pull them for you.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→Apple Machine Learning Research - How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering?, 10.2026→arXiv - How Much of a Harness Does a Strong Agent Need for Autonomous ML Engineering? (2609.40303), 30.09.2026
Original: https://wearecoded.com/en/articles/apple-harness-silen-agent.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news