The Allen Institute for AI has open-sourced AstaBrief-8B, a model that turns a research question and a set of retrieved literature excerpts into a fully cited report in a single pass. It already powers a new Fast mode inside Asta, the institute's agentic platform for scientific work, running alongside a Claude-powered Thinking mode that Ai2 is keeping for heavier jobs.
Key takeaways
- AstaBrief-8B ships under Apache 2.0 and is built on Qwen3-8B, with the supervised fine-tuning and preference datasets published alongside the weights.
- Across the full Asta pipeline, Fast mode averages 51.1 seconds per report against 178.5 seconds for the Claude-backed Thinking mode, about 3.5 times faster.
- The model wins 72 percent of judged report comparisons against Asta ScholarQA on the SQABench-CS2 test split, but scores 53.50 on DeepScholarBench where ScholarQA reaches 60.25.
How a one-pass pipeline cut report time
Ai2's existing report generator, ScholarQA, works in stages. It retrieves literature, summarizes and clusters the retrieved snippets, organizes the material into sections, then writes the report section by section. AstaBrief skips the summarization and clustering stages outright and writes the whole document at once from the query and the raw excerpts.
Ai2 says it expected to pay for that shortcut in quality and did not. The timing gap is the headline result: measured end to end, Fast mode returns a report in 51.1 seconds on average, against 178.5 seconds for Thinking mode.
Why Ai2 skipped reinforcement learning
The recipe deliberately avoids RL. The institute had already shown with DR Tulu that reinforcement learning can improve long-form report generation for open-weights models, but it settled on supervised fine-tuning followed by direct preference optimization instead, calling RL runs unstable, expensive and harder to debug.
That pushed the weight of the project onto data curation. Real Asta queries were filtered for length, language, scientific relevance and personal information, leaving a pool of 90,000 research questions. Report targets were generated through the ScholarQA pipeline using a mix of Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini and GPT-4.1, which yielded 47,000 usable fine-tuning examples. For the preference stage, GPT-4.1 and DeepSeek-R1 judged competing reports and only pairs where both judges agreed were kept, leaving roughly 6,000 examples. Ai2 reports 95 percent agreement between those judges and human preferences.
Of four statistical filters tested on the synthetic training reports, citation density β the share of statements carrying at least one citation β produced the clearest gains. More aggressive filtering and filter combinations added nothing meaningful.
Where it wins and where it does not
On the SQABench-CS2 test set, the model card puts AstaBrief-8B at an average of 87, against 83.7 for the fine-tuning-only checkpoint and 77.3 for the base Qwen3-8B. Citation recall moves furthest, from 64.6 on the base model to 78.2.
Against stronger systems the picture is mixed. AstaBrief wins 55 percent of judged comparisons against Asta ScholarQA on the dev split and 72 percent on test, beating DR-Tulu-8B's 36 and 54 percent. On DeepScholarBench it trails both, at 53.50 against 56.26 for DR-Tulu-8B and 60.25 for ScholarQA. In a separate 14-question human study, three researchers ranked DR-Tulu highest on overall preference, while two of the three put AstaBrief first on citation accuracy.
Ai2 attaches an unusual caveat. Most of the training and evaluation finished in 2025, and the comparison has not been rerun against current frontier models. The numbers, it says, are best read as evidence about the training and system design rather than about where this base model sits today.
What open weights change for research institutions
The licensing matters for the use case. Research questions often reveal unpublished or sensitive work, and open weights let an institution run report generation on its own hardware instead of sending queries to a proprietary API. Ai2 is shipping an example workflow for generating reports from a user's own PDFs.
Early adoption inside Asta is modest but sticky. Among 374 users who tried Fast mode, 29.1 percent used it on two or more days, generating 3.67 report threads on average. Twenty-three percent never switched back to Thinking mode, and another 18 percent alternated, using Fast mode for about 40 percent of their threads. Positive feedback ran at 84.2 percent, close to Thinking mode's 85.2 percent.
Outlook
Ai2 lists finer-grained preference learning, stronger combinations of retrieval-augmented generation and RL, multi-turn and multi-tool capability, and query decomposition as next steps. It also wants evaluations that go beyond checking whether a claim carries a citation to ask whether a model preserves the evidentiary scope of its sources β whether it quietly turns a sample-specific finding into a broad generalization. The work sits in the same push toward purpose-built research models as Intern-S2-Preview, and Ai2 says the data and attribution lessons will feed future versions of Olmo.
FAQ
Is AstaBrief-8B open source?
The weights are released under Apache 2.0 and are based on Qwen3-8B. Ai2 also published the fine-tuning and preference datasets used to train it, and says the model is intended for research and educational use in line with its Responsible Use Guidelines.
Does AstaBrief-8B beat the Claude-powered pipeline?
Not uniformly. It wins most judged head-to-head report comparisons against Asta ScholarQA on the SQABench-CS2 splits, but scores lower on DeepScholarBench and lost on overall preference in Ai2's small human study. The claim Ai2 makes is that it retains comparable quality while running several times faster.
What was it trained on?
It starts from Qwen3-8B, goes through supervised fine-tuning on 47,000 reports generated by proprietary models through the ScholarQA pipeline, then direct preference optimization on roughly 6,000 judged report pairs. The preference stage was run in Ai2's open-instruct framework on eight H100 GPUs.






