Thomson Reuters Launches Legal-Focused LLM After $40M Build
The information giant trained a domain-specific model on decades of legal content to power document review in its CoCounsel assistant.

Thomson Reuters debuts proprietary legal AI model
Thomson Reuters has launched Thomson, its first proprietary large language model designed specifically for legal work, the company announced Monday. The model will initially power Tabular Analysis, a document review feature within the CoCounsel Legal AI assistant that handles high-volume legal document processing.
The company invested approximately $40 million over two years in personnel and computing resources to develop the model, though the final training run cost just $450,000 due to efficiency gains. Rather than building from scratch, Thomson Reuters started with an open-weight foundation model and layered on its proprietary legal content, specialized training methods, and professional expertise.
CoCounsel will continue operating as a multi-model system, deploying Thomson where domain expertise provides advantages and using third-party frontier models for other tasks. Administrators can override the default and select alternative models for Tabular Analysis if needed.
Why it matters
Thomson Reuters' approach represents a middle path between relying entirely on general-purpose AI and building massive foundation models from scratch. For enterprises with deep domain expertise and proprietary content libraries, this strategy offers a blueprint for creating specialized AI capabilities without matching the multi-billion-dollar budgets of frontier labs. The legal industry's need for accuracy and citation verification makes it an ideal testing ground for domain-specific models that prioritize precision over breadth.
Training on 150 years of legal publishing
The development process involved realigning the base model with Thomson Reuters' values, pretraining on the company's content archive, and targeted post-training guided by legal professionals. The model learned to integrate with Thomson Reuters tools including Westlaw, which contains over 40,000 databases and more than 150 years of legal publishing and editorial curation.
Hundreds of subject-matter experts contributed to defining training objectives, creating sample legal questions, and evaluating responses through blind comparisons, according to Joel Hron, global head of artificial intelligence at Thomson Reuters.
The research team focused on continual learning techniques to add legal specialization without degrading the model's general capabilities. "If you simply take the open-source model without any of these additional steps, it won't know as much," said Jonathan Schwartz, head of foundational research at Thomson Reuters.
Performance and validation plans
Internal testing showed Thomson performing competitively with leading models when all had web-only access, according to Andrew Bean, a senior research scientist at the company. Performance improved to roughly equal or slightly better when connected to Thomson Reuters' proprietary content. Tests evaluated both answer completeness and citation accuracy.
These results have not yet undergone extensive independent validation. Thomson Reuters is sharing the model with legal experts and academic institutions for testing and plans to release a smaller open-weight version on Hugging Face under a noncommercial academic license. The company is also building a portal where outside developers can request API access for direct testing.
Only about 10 percent of Thomson Reuters' total information base has been incorporated so far. Future development will focus on converting the most valuable content and product usage patterns into better training signals rather than simply adding more material.
Hron acknowledged questions about whether Thomson Reuters can maintain pace with faster-moving AI laboratories. He argued that improvements in open models will provide stronger foundations for future versions while the company concentrates investment on professional applications. "I don't see owning an AI model that embodies the knowledge and expertise that TR possesses as something that's non-core to what we do," he said.
The details were first reported by SiliconANGLE.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call