Thomson Reuters Builds $40M Proprietary AI Model for Legal Work
The company specialized an open-weight foundation model with its own legal content, demonstrating a middle path between building from scratch and total reliance on frontier providers.
A corporate answer to AI sovereignty
Thomson Reuters has launched Thomson, a proprietary large language model tailored for legal work, offering the clearest blueprint yet for what AI sovereignty means at the enterprise level. The company invested approximately $40 million to develop the model after acquiring startup Safe Sign Technologies, according to details first reported by Forbes.
Rather than training a foundation model from scratch—a path requiring billions in capital—Thomson Reuters took what the company calls a "third way." It started with Alibaba's Qwen3.5-397B open-weight model, then specialized it using content from Westlaw, Practical Law, Checkpoint, and Reuters, combined with expert-generated training data.
The final three-week training run cost less than $450,000 in GPU expenses, though that figure excludes acquisition costs, staff, infrastructure, data preparation, and ongoing operations. The development team included no more than three dozen engineers and scientists, supported by hundreds of subject-matter experts who created more than 11,000 internal evaluation items.
Performance and limitations
According to Thomson Reuters' technical report, the Thomson-1.0-Large model's aggregate benchmark score exceeded GPT-5.4 and Claude Sonnet 5, though it trailed Claude Opus 4.8. The model showed weaknesses in coding, abstract reasoning, mathematics, and some robustness measures compared to leading general-purpose models.
In a blind study involving 35 attorney-editors and more than 3,000 comparisons, participants generally preferred the complete Thomson system over OpenAI and Anthropic systems for legal work. However, this compared full systems rather than isolated models—Thomson had access to Westlaw and other proprietary tools, while competitors used web search. The study was conducted by Thomson Reuters and has not been independently replicated.
CEO Steve Hasker noted the company has used less than 10 percent of its legal content to date, suggesting substantial room for improvement.
Why it matters
Thomson Reuters demonstrates that companies with proprietary, rights-cleared data and domain experts can reduce strategic dependence on frontier AI providers without matching their capital expenditure. This matters for data-rich institutions—insurers, healthcare systems, financial firms—evaluating whether to build specialized models or rely entirely on external providers. The approach offers greater control over intellectual property and potentially lower long-term inference costs, though it requires substantial upfront investment and ongoing maintenance. Thomson Reuters is exploring deployments that would place the model inside customer cloud environments, allowing clients to use AI with their sensitive documents without sending data to third-party providers.
The sovereignty spectrum
Thomson Reuters has not eliminated dependency—it rents compute, relies on GPU hardware, and started from an open-weight foundation it didn't create. But it exercises control over the layers that differentiate its business: proprietary content, expert judgment, evaluation standards, and deployment decisions.
The company established a Frontier AI Research Lab with Imperial College London and used partners including DatologyAI and Lambda. It rented and ring-fenced compute in the cloud rather than building its own data center, and says it changed the root model multiple times as open-weight options improved.
Before pursuing a similar strategy, companies should assess whether they possess data competitors cannot obtain, have experts who can define correctness, can identify high-value workflows that justify the investment, and are prepared to maintain and migrate the model as technology evolves.
These details were first reported by John Sviokla in Forbes.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
