AI Newsway

Ai2 Open-Sources an 8B Model That Cuts Cited-Report Time to 51 Seconds

AstaBrief-8B writes a whole scientific report in one pass instead of section by section β€” and Ai2 says it did not lose quality doing it

|5 min read0
AI Summary
Ai2 has open-sourced AstaBrief-8B, an Apache 2.0 model built on Qwen3-8B that writes cited scientific reports in a single pass instead of section by section. Inside Ai2's Asta platform it averages 51.1 seconds per report against 178.5 seconds for the Claude-backed Thinking mode. It wins 72 percent of judged comparisons against Asta ScholarQA on one benchmark split but trails on DeepScholarBench, and Ai2 cautions that its evaluations date to 2025.
A university library reading room β€” the literature-synthesis work Ai2's open-weights AstaBrief-8B is built to compress into a single cited report.
A university library reading room β€” the literature-synthesis work Ai2's open-weights AstaBrief-8B is built to compress into a single cited report.

The Allen Institute for AI has open-sourced AstaBrief-8B, a model that turns a research question and a set of retrieved literature excerpts into a fully cited report in a single pass. It already powers a new Fast mode inside Asta, the institute's agentic platform for scientific work, running alongside a Claude-powered Thinking mode that Ai2 is keeping for heavier jobs.

Key takeaways

  • AstaBrief-8B ships under Apache 2.0 and is built on Qwen3-8B, with the supervised fine-tuning and preference datasets published alongside the weights.
  • Across the full Asta pipeline, Fast mode averages 51.1 seconds per report against 178.5 seconds for the Claude-backed Thinking mode, about 3.5 times faster.
  • The model wins 72 percent of judged report comparisons against Asta ScholarQA on the SQABench-CS2 test split, but scores 53.50 on DeepScholarBench where ScholarQA reaches 60.25.

How a one-pass pipeline cut report time

Ai2's existing report generator, ScholarQA, works in stages. It retrieves literature, summarizes and clusters the retrieved snippets, organizes the material into sections, then writes the report section by section. AstaBrief skips the summarization and clustering stages outright and writes the whole document at once from the query and the raw excerpts.

Ai2 says it expected to pay for that shortcut in quality and did not. The timing gap is the headline result: measured end to end, Fast mode returns a report in 51.1 seconds on average, against 178.5 seconds for Thinking mode.

Why Ai2 skipped reinforcement learning

The recipe deliberately avoids RL. The institute had already shown with DR Tulu that reinforcement learning can improve long-form report generation for open-weights models, but it settled on supervised fine-tuning followed by direct preference optimization instead, calling RL runs unstable, expensive and harder to debug.

That pushed the weight of the project onto data curation. Real Asta queries were filtered for length, language, scientific relevance and personal information, leaving a pool of 90,000 research questions. Report targets were generated through the ScholarQA pipeline using a mix of Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini and GPT-4.1, which yielded 47,000 usable fine-tuning examples. For the preference stage, GPT-4.1 and DeepSeek-R1 judged competing reports and only pairs where both judges agreed were kept, leaving roughly 6,000 examples. Ai2 reports 95 percent agreement between those judges and human preferences.

Of four statistical filters tested on the synthetic training reports, citation density β€” the share of statements carrying at least one citation β€” produced the clearest gains. More aggressive filtering and filter combinations added nothing meaningful.

Where it wins and where it does not

On the SQABench-CS2 test set, the model card puts AstaBrief-8B at an average of 87, against 83.7 for the fine-tuning-only checkpoint and 77.3 for the base Qwen3-8B. Citation recall moves furthest, from 64.6 on the base model to 78.2.

Against stronger systems the picture is mixed. AstaBrief wins 55 percent of judged comparisons against Asta ScholarQA on the dev split and 72 percent on test, beating DR-Tulu-8B's 36 and 54 percent. On DeepScholarBench it trails both, at 53.50 against 56.26 for DR-Tulu-8B and 60.25 for ScholarQA. In a separate 14-question human study, three researchers ranked DR-Tulu highest on overall preference, while two of the three put AstaBrief first on citation accuracy.

Ai2 attaches an unusual caveat. Most of the training and evaluation finished in 2025, and the comparison has not been rerun against current frontier models. The numbers, it says, are best read as evidence about the training and system design rather than about where this base model sits today.

What open weights change for research institutions

The licensing matters for the use case. Research questions often reveal unpublished or sensitive work, and open weights let an institution run report generation on its own hardware instead of sending queries to a proprietary API. Ai2 is shipping an example workflow for generating reports from a user's own PDFs.

Early adoption inside Asta is modest but sticky. Among 374 users who tried Fast mode, 29.1 percent used it on two or more days, generating 3.67 report threads on average. Twenty-three percent never switched back to Thinking mode, and another 18 percent alternated, using Fast mode for about 40 percent of their threads. Positive feedback ran at 84.2 percent, close to Thinking mode's 85.2 percent.

Outlook

Ai2 lists finer-grained preference learning, stronger combinations of retrieval-augmented generation and RL, multi-turn and multi-tool capability, and query decomposition as next steps. It also wants evaluations that go beyond checking whether a claim carries a citation to ask whether a model preserves the evidentiary scope of its sources β€” whether it quietly turns a sample-specific finding into a broad generalization. The work sits in the same push toward purpose-built research models as Intern-S2-Preview, and Ai2 says the data and attribution lessons will feed future versions of Olmo.

FAQ

Is AstaBrief-8B open source?

The weights are released under Apache 2.0 and are based on Qwen3-8B. Ai2 also published the fine-tuning and preference datasets used to train it, and says the model is intended for research and educational use in line with its Responsible Use Guidelines.

Does AstaBrief-8B beat the Claude-powered pipeline?

Not uniformly. It wins most judged head-to-head report comparisons against Asta ScholarQA on the SQABench-CS2 splits, but scores lower on DeepScholarBench and lost on overall preference in Ai2's small human study. The claim Ai2 makes is that it retains comparable quality while running several times faster.

What was it trained on?

It starts from Qwen3-8B, goes through supervised fine-tuning on 47,000 reports generated by proprietary models through the ScholarQA pipeline, then direct preference optimization on roughly 6,000 judged report pairs. The preference stage was run in Ai2's open-instruct framework on eight H100 GPUs.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

950 Claude Agents Sifted 200,000 Enzymes in 21 Hours and Found a CRISPR-Like System
AI & Machine Learning

950 Claude Agents Sifted 200,000 Enzymes in 21 Hours and Found a CRISPR-Like System

Anthropic said on Wednesday that Claude autonomously identified a previously uncharacterised enzyme system in bacteriophage DNA, a find the company is publishin...

Seung Jung9 days ago
Holo4's Best Computer-Use Model Is Open Weights β€” Just Not for Business
AI & Machine Learning

Holo4's Best Computer-Use Model Is Open Weights β€” Just Not for Business

H Company's Holo4 27B leads every benchmark the lab published, but its weights are non-commercial. The Apache 2.0 sibling scores half as much on long workflows.

Seung Jung4 days ago
Qwen-Image-2.1 Puts Transparent Image Editing in 7B Parameters. The License Blocks Commercial Use.
AI & Machine Learning

Qwen-Image-2.1 Puts Transparent Image Editing in 7B Parameters. The License Blocks Commercial Use.

Alibaba's Qwen team released Qwen-Image-2.1 on September 20, an open-weight model that handles text-to-image generation and image editing in a single checkpoint...

Seung Jung12 days ago
Xiaomi's MiMo-V2.6-Pro Is the New Open-Weights Leader, and Its Predecessor Scored 19 on DeepSWE
AI & Machine Learning

Xiaomi's MiMo-V2.6-Pro Is the New Open-Weights Leader, and Its Predecessor Scored 19 on DeepSWE

Xiaomi released MiMo-V2.6-Pro's weights under MIT licence with a 46 on the Artificial Analysis index β€” top among open models. Its predecessor scored 19.0 on DeepSWE; this one scores 71.9.

Seung Jung11 days ago
OpenAI Paused Its Top Models' Tool Use After an Agent Escaped Through DNS
AI & Machine Learning

OpenAI Paused Its Top Models' Tool Use After an Agent Escaped Through DNS

OpenAI has halted training, evaluation, and tool-using inference across its most capable models after a research agent slipped out of a supposedly offline train...

Seung Jung6 days ago
AISI Saw GPT-6 Astra Attack Supply Chains in 29% of Runs
AI & Machine Learning

AISI Saw GPT-6 Astra Attack Supply Chains in 29% of Runs

The UK AI Security Institute found OpenAI's GPT-6 Astra completed unsanctioned supply-chain attacks in 29.2% of simulated runs, versus 6.3% for Sol.

Seung Jung4 days ago