we_are_coded.by CODE · The world, decoded
БГ
Ai2

Ai2 open-sources AstaBrief 8B, the model behind Asta's Fast mode: a cited report about 3.5 times faster than the Claude mode, per Ai2

Ai2 · event date: 2 October 2026Frontier

Ai2 has published the model and its training data. From a research question and excerpts from the literature, AstaBrief writes a cited report, and in Asta it runs as Fast mode next to the Claude-powered Thinking mode. Ai2 itself says the training and evaluation were done mostly in 2025 and that the full evaluation has not been rerun against today's leading models.

In short
  • AstaBrief 8B starts from Qwen3-8B and writes the whole report in one go, not section by section. Per Ai2, that is how a report in Asta comes out about 3.5 times faster than in the Claude-powered Thinking mode.
  • The examples for the main training stage were written by closed models such as Claude, o3 and GPT-4.1. The most useful was a simple filter: how many claims in a report carry a citation.
  • The human study is small: 14 questions, three researchers. DR-Tulu wins overall preference; on citation accuracy two of the three pick AstaBrief.
Checked on5 October 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

51 seconds for one cited report. That is the average in Asta, per Ai2, when its small model does the writing instead of the Claude mode, which needs almost three minutes. The post is more interesting than the number, though.

The facts: on 2 October 2026 Ai2 published AstaBrief 8B, a model that turns a research question and excerpts from scientific literature into a cited report. The model and the training data are open. In the Generate a report feature of the Asta platform it runs as Fast mode, next to Thinking mode, which is Claude-powered. AstaBrief starts from Qwen3-8B and writes the whole report in one pass, skipping the intermediate steps of summarising and clustering snippets and not writing section by section. Per Ai2, the average time across the full Asta pipeline is 51.1 seconds per report against 178.5 for Thinking, about 3.5 times faster. The starting pool is 90 thousand research queries from real Asta users, left after filtering for quality, relevance and personal information. For SFT, the target reports were generated with Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini and GPT-4.1, and after filtering 47 thousand examples remain. For DPO there are about 6 thousand pairs of reports where two judge models, GPT-4.1 and DeepSeek-R1, picked the same winner. The main test is SQABench-CS2, 200 computer science questions; per Ai2, AstaBrief is competitive with the Claude pipeline and with DR Tulu on several measures of answer and citation quality. In a separate study with three researchers and 14 questions, DR-Tulu wins overall preference, while two of the three prefer AstaBrief on citation accuracy. Ai2 notes that most of the training and evaluation was completed in 2025 and that the full evaluation has not been rerun against today's leading models. Of the 374 Asta users who tried Fast mode, 23% never went back to Thinking. With the weights, Ai2 releases an example workflow for reports from your own PDFs.

Look at who it learns from. The example reports for the main training stage were written by closed models: Claude, o3, GPT-4.1. The small open model is a student of the big ones. The model and the data are open. Its teachers are not. Ai2 says so openly, and you should read that next to the speed.

Speed is the easy part. Correctness is the hard one. The post says itself that a citation next to a sentence does not make the sentence true: a model can point at the exact study and still claim more than the study shows. A finding about one sample becomes a rule for everyone, past tense becomes present, and the description of a result turns into advice on what clinicians, policymakers or researchers should do.

A citation next to a sentence does not make it true.

Ai2 admits the hole in its own yardstick. They mainly measured relevance, coverage and whether the citations support the claims. Whether the claim stayed within the scope of its source, they do not measure yet and list it among the next tasks.

The most useful thing in the post is modest. They tried four filters for weak training examples and a simple one helped most: drop the reports where few claims carry a citation. More aggressive filtering and combinations did not add much. Ai2 states it as a lesson: specialisation is not necessarily a matter of more scientific text; a lot depends on whether the training examples show the behaviour.

The comparisons are Ai2's and were made mostly in 2025. In the study with the researchers another model wins overall preference, and there are only 14 questions.

There is a real use, too: an institution whose question is sensitive or still unpublished can run AstaBrief on its own machines, behind its own firewall, without someone else's API. Run it on your own PDFs and open the citations one by one.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→Ai2 - Open-sourcing AstaBrief, the fast report-generation model in Asta, 02.10.2026
Original: https://wearecoded.com/en/articles/ai2-astabrief-8b-otvoren-model.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news