Ai2 has published the model and its training data. From a research question and excerpts from the literature, AstaBrief writes a cited report, and in Asta it runs as Fast mode next to the Claude-powered Thinking mode. Ai2 itself says the training and evaluation were done mostly in 2025 and that the full evaluation has not been rerun against today's leading models.
- AstaBrief 8B starts from Qwen3-8B and writes the whole report in one go, not section by section. Per Ai2, that is how a report in Asta comes out about 3.5 times faster than in the Claude-powered Thinking mode.
- The examples for the main training stage were written by closed models such as Claude, o3 and GPT-4.1. The most useful was a simple filter: how many claims in a report carry a citation.
- The human study is small: 14 questions, three researchers. DR-Tulu wins overall preference; on citation accuracy two of the three pick AstaBrief.
51 seconds for one cited report. That is the average in Asta, per Ai2, when its small model does the writing instead of the Claude mode, which needs almost three minutes. The post is more interesting than the number, though.
Look at who it learns from. The example reports for the main training stage were written by closed models: Claude, o3, GPT-4.1. The small open model is a student of the big ones. The model and the data are open. Its teachers are not. Ai2 says so openly, and you should read that next to the speed.
Speed is the easy part. Correctness is the hard one. The post says itself that a citation next to a sentence does not make the sentence true: a model can point at the exact study and still claim more than the study shows. A finding about one sample becomes a rule for everyone, past tense becomes present, and the description of a result turns into advice on what clinicians, policymakers or researchers should do.
Ai2 admits the hole in its own yardstick. They mainly measured relevance, coverage and whether the citations support the claims. Whether the claim stayed within the scope of its source, they do not measure yet and list it among the next tasks.
The most useful thing in the post is modest. They tried four filters for weak training examples and a simple one helped most: drop the reports where few claims carry a citation. More aggressive filtering and combinations did not add much. Ai2 states it as a lesson: specialisation is not necessarily a matter of more scientific text; a lot depends on whether the training examples show the behaviour.
The comparisons are Ai2's and were made mostly in 2025. In the study with the researchers another model wins overall preference, and there are only 14 questions.
There is a real use, too: an institution whose question is sensitive or still unpublished can run AstaBrief on its own machines, behind its own firewall, without someone else's API. Run it on your own PDFs and open the citations one by one.