Thomson Reuters has launched Thomson, its first in-house large language model. The company announced it Monday from Toronto. The claim underneath the launch is unusual for a legal and tax publisher: it now owns the model, not only the archive it was trained on.
The budget is the headline number. Thomson Reuters put $40 million into training, a figure that covers both talent and compute. Frontier labs have routinely spent billions and several years of infrastructure work to reach comparable capability.
The company did not start from nothing. It began with an open-source foundation, then layered mid-training and post-training on decades of proprietary material drawn from Westlaw, Practical Law, Checkpoint and Reuters. Hundreds of subject matter experts were involved, from the design of training objectives through to final evaluation.
Specialization instead of scale
Chief technology officer Joel Hron argued that the industry has treated scale as the answer for years, and said a strong foundation specialized deeply for particular work can produce intelligence that is highly capable, far more efficient and entirely under the company's control.
Hron has been blunter elsewhere about the motive. He told Legal IT Insider the model was built first for Thomson Reuters itself. The aim was to reduce dependence on someone else's roadmap as access to powerful models becomes commoditized.
He also drew a line around what is actually defensible. The open-source base is the starting point, not the moat. The hard part is the data, the training methodology, the preference data and the evaluation infrastructure needed to tell whether a model is improving at professional work.
The evidence so far
Chief executive Steve Hasker said early evaluations place Thomson on par with the latest frontier models across a range of tasks. The company reports a meaningful uplift over its base model in instruction following, and a larger uplift in navigating dense, domain-specific material.
Outside academics have begun testing it. Jonathan Choi of Washington University School of Law ran corporate tax questions against Thomson, ChatGPT and Claude. All three answered correctly, but he preferred Thomson's responses and singled out its links to treatises as more transparent for legal work.
Samuel Dahan, who directs labs at Queen's and Cornell, found citation quality generally competitive with leading frontier models, even on Canadian employment law without a Canada-specific configuration.
Both assessments are early and narrow. Neither is a published benchmark, and Thomson Reuters chose which reviewers received access first. The company says it will widen external evaluation over the coming weeks, and is releasing a small open-weight version on Hugging Face for academic and non-commercial use.
Where it actually ships
The first deployment is Tabular Analysis inside CoCounsel Legal. That is high-volume, structured document review, the workload where a purpose-built model's advantage shows up soonest.
CoCounsel stays multi-model by design. Thomson Reuters will continue routing to OpenAI and Anthropic models where those perform better, and to its own model where it wins. The company frames this as orchestration rather than replacement.
Only a fraction of the corpus has been touched. Training so far has used less than 10% of Thomson Reuters content. Hron said the next step is not simply feeding the model more data, but identifying which content actually improves performance.
A strategy years in the making
The launch traces back to the 2024 acquisition of UK startup Safe Sign Technologies. At the time that deal read as a bolt-on for CoCounsel. In hindsight it looks like the opening move.
Hron confirmed the company is open to licensing Thomson directly to law firms, corporate legal departments and other organizations that want more control over their own AI stack. That would put Thomson Reuters in the business of supplying intelligence, not just applications built on top of it. For a company whose rivals are legal AI startups renting frontier capacity by the token, owning the model rewrites the cost structure beneath every product it sells.






